maxbbraun / maxbbraun/llama4micro
Optimize inference speed ⚡️
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 561
- Forks
- 37
- PR merge metrics
- No merged PRs in 30d
Description
Experimenting with compiler options in branch fast-opts.
Switching from -Os to -O3 seems to have significant impact on tokens per second. (-Ofast doesn't noticeably add on top.)
->>> Averaged 2.60 tokens/s
+>>> Averaged 3.79 tokens/s
Unfortunately, something about this seems to break the camera input or TPU inference and I haven't debugged that yet.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the fast-opts branch and compare its compiler-option change from -Os to -O3. Reproduce the token-per-second benchmark, then investigate the breakage affecting camera input or TPU inference. Done means the speed improvement is retained without breaking either inference path.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- embedded-iot, machine-learning, performance
- Issue type
- Refactor
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100