AlibabaResearch

AlibabaResearch/flash-llm

View on GitHub

Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Stars
247
Forks
24
Open beginner issues
0
Indexed issues
7
Dominant language
Cuda
License
Apache-2.0
Last GitHub push
Sep 22, 2023
Latest indexed
Sep 13, 2026
Contributing guide
No contributing guide
Code of conduct
No code of conduct
Beginner labels
No beginner labels indexed
PR merge metrics
No merged PRs in 30d
7 open issues indexed Loading issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.