deepspeedai / deepspeedai/DeepSpeed
[REQUEST] Faster quantization with Smoothquant
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Smoothquant is a recent innovation that seems to enable state of the art faster inference for INT8 (and probably any precision)
https://github.com/mit-han-lab/smoothquant
SmoothQuant migrates part of the quantization difficulties from activation to weights, which smooths out the systematic outliers in activation, making both weights and activations easy to quantize.
It would be impactful if deepspeed could integrate this optimization and achieve faster inference.
@Guangxuan-Xiao
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no DeepSpeed files or tests. Start by reviewing the linked SmoothQuant work and locating DeepSpeed’s inference quantization entry point; define the integration scope and a benchmark or test that demonstrates faster INT8 inference before implementation.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100