tensorflow / tensorflow/model-optimization
about the Quantize layer when trans model
@sngyhan is already working on this.
Since Aug 1, 2022.
- Dominant language
- Python
- Stars
- 1.6k
- Forks
- 349
- Avg merge
- 3d 2h
- Merged PRs (30d)
- 1
Description
I have two model, which is mobilenetv1 for classification.
the first model, it's download from google: https://storage.googleapis.com/download.tensorflow.org/models/tflite/mobilenet_v1_224_android_quant_2017_11_08.zip
the second model, it's ctreat by myself, I make its layers same to first model, it is train by keras and Post-training quantization(PTQ) to get tflite model which input/output are 'uint8', but my model have two layer 'Quantize', which is in my model's head and tail. just like image below.
the first model run on my npu is about 8ms, but the second model is 30ms, what happen? it just diff only two 'Quantize' layer. so, what can I do, I follow the sample 'https://tensorflow.google.cn/lite/performance/post_training_integer_quant' to train my model, but it is slow than official model, and a little diff from official model, please help, some suggestions, or some other guide or sample code to get the 'uint8' model.


Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.