huggingface / huggingface/optimum-executorch

quantization for own model

Open
#202 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
141
Forks
48
Avg merge
26m
Merged PRs (30d)
2

Description

Can I perform the quantization using the "optimum" for my own executorch FP16 model, which needs to be quantized as q4fp16? Here, my model includes mamba blocks.
If yes, then can you guide with step by step process?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.