huggingface / huggingface/swift-chat
Generation speed issue
- Dominant language
- Swift
- Stars
- 597
- Forks
- 49
- PR merge metrics
- No merged PRs in 30d
Description
I load llama2 model like example successfully but the speed to generate text is really slow.

[1] I'm not sure it use mps to accelerate generation.
How to confirm it?
[2] Is there a smaller LLM than 7B?
Here is my env
- Macbook Air / M2 / 16GB / Sonoma 14.5
- Xcode 15.4
- ckpt: coreml-projects/Llama-2-7b-chat-coreml
Contributor guide
No contributing guide indexed for this repository
Research direction
Reproduce the slow generation on the stated MacBook Air M2 environment using the llama2 example and the listed coreml-projects/Llama-2-7b-chat-coreml checkpoint. Determine whether MPS acceleration is active and document the observed generation speed; the issue does not mention any source files or tests, and completion criteria are not defined beyond answering the two questions.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- swift
- Domain
- ai, desktop, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100