huggingface / huggingface/candle
Add support for the quantized llama 3.2 models
Open
- Dominant language
- Rust
- Stars
- 21k
- Forks
- 1.8k
- Avg merge
- 16h 42m
- Merged PRs (30d)
- 25
Description
The new Llama3.2 models have just been uploaded to HF :
https://huggingface.co/models?other=arxiv:2405.16406
models showing much faster inference and smaller size. Would be great to add them
Thanks
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no files, tests, or entry points. Start by reviewing the quantized Llama 3.2 models in the linked Hugging Face listing and locating Candle's existing model-support entry points; done means Candle supports inference for these models with their smaller size and faster inference.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100