[Upstream]: vLLM inference with AR kernels
Open
high priority
- Dominant language
- Python
- Stars
- 1.6k
- Forks
- 175
- Avg merge
- 1d 18h
- Merged PRs (30d)
- 99
Description
### Feature Description
Load the AR quantized model and inference with AR kernels.
### Motivation and Use Case
Support AR in the mainstream inference framework
### Alternatives Considered
_No response_
### Definition of Done
_No response_
### Additional Context
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.