deepspeedai / deepspeedai/DeepSpeed

can deepspeed fast-gen support int8 weightonly inference

Open
#4,786 0 comments 0 reactions 1 assignee View on GitHub

@HeyangQin is already working on this.

Since Dec 15, 2023.

enhancement inference
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

I test the fastgen on codellama-34b a100, the speed is as:
tensorrt llm 1a100 19token/s, fast-gen 21token/s
the speed of 2a100 on fast-gen is 36token/s

however the speed of tensorrt llm 1a100 is 46token/s on int4 quant, 30token/s on int8.

int4/int8 quantization is the best inference speedup strategy, so can fast-gen support int4/int8 weightonly.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.