alibaba / alibaba/TinyNeuralNetwork

Support W16A16 quantization

Open
#391 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
879
Forks
134
PR merge metrics
No merged PRs in 30d

Description

Hi Authors,

Some of our models need to be quantized as int16 (w16a16) to achieve better accuracy. Looks tinynn currently does not support that. Please kindly evaluate the fessibility of implementation. Big thanks.

Contributor guide

Open the contributing guide

Research direction

No files, tests, or entry points are named. Start by reviewing TinyNeuralNetwork's existing quantization support and determine what would be required for W16A16, then define completion as successful int16 weight and activation quantization for representative models with coverage in the relevant tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.