huggingface / huggingface/huggingface-inference-toolkit
OOM error
Open
- Dominant language
- Python
- Stars
- 97
- Forks
- 28
- Avg merge
- 6d 23m
- Merged PRs (30d)
- 4
Description
Is there any env var or flag i can set to allocate only 80% of the gpu's memory? i get occasional OOMs because the container is keeping the VRAM usage at high 90% +.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.