How to run the turbo version on a 24G graphics card?
Open
- Dominant language
- Python
- Stars
- 508
- Forks
- 35
- PR merge metrics
- No merged PRs in 30d
Description
I have tested both the SGLang and Diffusers versions on a 24G RTX 4090 graphics card. While the model can be successfully loaded, an out-of-memory (OOM) error occurs during the inference process.
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue mentions the SGLang and Diffusers versions, a 24G RTX 4090, and an OOM during inference, but names no files or tests. Start by reproducing the turbo-version inference setup on that hardware and comparing the memory behavior of both versions; done means documenting a confirmed way to run it or the applicable memory limitation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100