intel / intel/auto-round

[feat] For FP8 models, keep layers specified in `ignore_layers` in their original FP8 format

Open
#1,284 10 comments 0 reactions 1 assignee Claimed by @yiliu30 View on GitHub
good first issue
Dominant language
Python
Stars
1.6k
Forks
175
Avg merge
1d 18h
Merged PRs (30d)
99

Description

For Deepseek v32 and similar FP8 models, it is preferable to keep layers specified in `ignore_layers` (such as indexer or attn) in their original FP8 format, rather than dequantizing them to BF16.

Expected behavior:

```
AR_LOG_LEVEL=TRACE auto_round --model /models/Qwen3-8B-FP8 --ignore_layers "attn"
```
All layers within attention (matching "attn") should remain in FP8 format, not be dequantized to BF16 or float.

Depends on https://github.com/intel/auto-round/issues/1283

cc @wenhuach21 @thuang6 @xin3he

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.