AlibabaResearch/flash-llm
View on GitHubFlash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
- Stars
- 247
- Forks
- 24
- Open beginner issues
- 0
- Indexed issues
- 7
- Dominant language
- Cuda
- License
- Apache-2.0
- Last GitHub push
- Sep 22, 2023
- Latest indexed
- Sep 13, 2026
- Contributing guide
- No contributing guide
- Code of conduct
- No code of conduct
- Beginner labels
- No beginner labels indexed
- PR merge metrics
- No merged PRs in 30d
-
AlibabaResearch/flash-llm#8 · 5 comments · 0 reactions · 0 assignees ·
-
AlibabaResearch/flash-llm#9 · 3 comments · 0 reactions · 0 assignees ·
-
smaller OPT? Open
AlibabaResearch/flash-llm#10 · 0 comments · 0 reactions · 0 assignees ·
-
AlibabaResearch/flash-llm#12 · 0 comments · 0 reactions · 0 assignees ·
-
AlibabaResearch/flash-llm#13 · 1 comment · 0 reactions · 0 assignees ·
-
AlibabaResearch/flash-llm#14 · 0 comments · 0 reactions · 0 assignees ·
-
AlibabaResearch/flash-llm#15 · 0 comments · 0 reactions · 0 assignees ·