huggingface / huggingface/candle

Add support for the quantized llama 3.2 models

Open
#2,573 0 comments 1 reaction 0 assignees View on GitHub
Dominant language
Rust
Stars
21k
Forks
1.8k
Avg merge
16h 42m
Merged PRs (30d)
25

Description

The new Llama3.2 models have just been uploaded to HF :
https://huggingface.co/models?other=arxiv:2405.16406

models showing much faster inference and smaller size. Would be great to add them

Thanks

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or entry points. Start by reviewing the quantized Llama 3.2 models in the linked Hugging Face listing and locating Candle's existing model-support entry points; done means Candle supports inference for these models with their smaller size and faster inference.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.