intel / intel/auto-round

[Feature]: support FP8+WOQ mix datatypes inference in transformers and vllm

Open
#2,225 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
1.6k
Forks
175
Avg merge
1d 18h
Merged PRs (30d)
99

Description

### Feature Description

~

### Motivation and Use Case

~

### Alternatives Considered

_No response_

### Definition of Done

_No response_

### Additional Context

_No response_

Contributor guide

Open the contributing guide

Research direction

The issue provides no files, tests, entry points, or definition of done. Start by locating the existing datatype inference paths for Transformers and vLLM, then clarify the expected FP8+WOQ combinations and validation criteria before implementation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.