alibaba / alibaba/rtp-llm

what's the decode token speed of 7b qwen gptq int4 model on v100?

Open
#131 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.3k
Forks
275
Avg merge
3d 17h
Merged PRs (30d)
33

Description

how about the compre with vllm? thx!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by identifying the repository entry points for running a 7B Qwen GPTQ INT4 model on a V100, then determine how decode-token speed is measured. Compare the result with vLLM under equivalent conditions and document the measurements and test setup; the issue is complete when both speeds are reported clearly.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, performance
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.