alibaba / alibaba/TinyNeuralNetwork
Support W16A16 quantization
Open
enhancement
- Dominant language
- Python
- Stars
- 879
- Forks
- 134
- PR merge metrics
- No merged PRs in 30d
Description
Hi Authors,
Some of our models need to be quantized as int16 (w16a16) to achieve better accuracy. Looks tinynn currently does not support that. Please kindly evaluate the fessibility of implementation. Big thanks.
Contributor guide
Research direction
No files, tests, or entry points are named. Start by reviewing TinyNeuralNetwork's existing quantization support and determine what would be required for W16A16, then define completion as successful int16 weight and activation quantization for representative models with coverage in the relevant tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100