vllm-project / vllm-project/aibrix
One question about "GPU Hardware Failure Detection"
Open
area/distributed
kind/support
- Dominant language
- Go
- Stars
- 5.1k
- Forks
- 694
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 98
Description

Inside one ray-cluster, if one GPU(ray work node) have hardware issue, how to recover this ray-cluster?
-Just replace this worker GPU?
-Replace this ray-cluster? Set this cluster Failure.
Could you provide any details about "GPU Hardware Failure Detection: Proactive detection of GPU hardware issues."
Contributor guide
Assessment
This issue has not been assessed yet.