huggingface / huggingface/candle
Reducing VRAM consumption
Open
- Dominant language
- Rust
- Stars
- 21k
- Forks
- 1.8k
- Avg merge
- 16h 42m
- Merged PRs (30d)
- 25
Description
[This article](https://huggingface.co/blog/lyogavin/airllm) suggest come optimization tweaks, which allows to run big models on potato GPUs. Can some of this tweaks be implemented in Candle?
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading the linked AirLLM article and comparing its optimization tweaks with Candle's current model execution and memory behavior. Identify which specific changes are applicable, then define completion by implementing selected tweaks and measuring reduced VRAM consumption on representative models.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100