huggingface / huggingface/candle

Reducing VRAM consumption

Open
#1,679 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
21k
Forks
1.8k
Avg merge
16h 42m
Merged PRs (30d)
25

Description

[This article](https://huggingface.co/blog/lyogavin/airllm) suggest come optimization tweaks, which allows to run big models on potato GPUs. Can some of this tweaks be implemented in Candle?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading the linked AirLLM article and comparing its optimization tweaks with Candle's current model execution and memory behavior. Identify which specific changes are applicable, then define completion by implementing selected tweaks and measuring reduced VRAM consumption on representative models.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
machine-learning, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.