huggingface / huggingface/swift-chat

Generation speed issue

Open
#26 2 comments 1 reaction 0 assignees View on GitHub
Dominant language
Swift
Stars
597
Forks
49
PR merge metrics
No merged PRs in 30d

Description

I load llama2 model like example successfully but the speed to generate text is really slow.

![image](https://github.com/user-attachments/assets/f3275132-9cce-4ac5-b928-803c6e66147f)

[1] I'm not sure it use mps to accelerate generation.
How to confirm it?
[2] Is there a smaller LLM than 7B?

Here is my env
- Macbook Air / M2 / 16GB / Sonoma 14.5
- Xcode 15.4
- ckpt: coreml-projects/Llama-2-7b-chat-coreml

Contributor guide

No contributing guide indexed for this repository

Research direction

Reproduce the slow generation on the stated MacBook Air M2 environment using the llama2 example and the listed coreml-projects/Llama-2-7b-chat-coreml checkpoint. Determine whether MPS acceleration is active and document the observed generation speed; the issue does not mention any source files or tests, and completion criteria are not defined beyond answering the two questions.

Written by the indexing model from the issue text.

Assessment

Tech stack
swift
Domain
ai, desktop, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.