huggingface / huggingface/optimum-executorch
quantization for own model
Open
- Dominant language
- Python
- Stars
- 141
- Forks
- 48
- Avg merge
- 26m
- Merged PRs (30d)
- 2
Description
Can I perform the quantization using the "optimum" for my own executorch FP16 model, which needs to be quantized as q4fp16? Here, my model includes mamba blocks.
If yes, then can you guide with step by step process?
Contributor guide
Assessment
This issue has not been assessed yet.