InternLM / InternLM/lmdeploy

[Feature] W8A8 support for turbomind engine

Open
#2,962 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
8.1k
Forks
748
Avg merge
6d 2h
Merged PRs (30d)
54

Description

### Motivation

For now, W8A8 quantization is only supported with pytorch engine. But the turbomind engine is much more performant. Do you have any plan to support W8A8 for turbomind engine.

Also, only SmoothQuant is supported as in [here](https://github.com/InternLM/lmdeploy/blob/main/docs/en/quantization/w8a8.md). I think FP8 quantization is also a feature to consider.

Thanks!

### Related resources

_No response_

### Additional context

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.