vllm-project / vllm-project/aibrix

One question about "GPU Hardware Failure Detection"

Open
#1,129 2 comments 0 reactions 1 assignee Claimed by @Jeffwan View on GitHub
area/distributed kind/support
Dominant language
Go
Stars
5.1k
Forks
694
Avg merge
1d 19h
Merged PRs (30d)
98

Description

![Image](https://github.com/user-attachments/assets/b3e1237a-b73c-43ec-bf4e-3ecdd8af0740)

Inside one ray-cluster, if one GPU(ray work node) have hardware issue, how to recover this ray-cluster?
-Just replace this worker GPU?
-Replace this ray-cluster? Set this cluster Failure.

Could you provide any details about "GPU Hardware Failure Detection: Proactive detection of GPU hardware issues."

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.