Publish VLLM and SGLANG configs where possible / 尽可能同时发布 vLLM 和 SGLang 配置
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 32/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Quiet
- Tech stack
- python
- Domain
- machine-learning, performance
Research direction
Start by locating how benchmark configurations and results are defined for the existing vLLM or SGLang runs. Compare the available model and hardware cases, then determine what is needed to publish both frameworks and how settings such as disable_prefix_cache=True should be handled. Done means the project has an agreed, repeatable policy and corresponding benchmark coverage.
Written by the indexing model from the issue text.
Description
All the benchmarks have either VLLM or sglang but not both. I assume this is to avoid twitter fights. I think a lot of people at labs like me are on a fork of one and spend some energy trying to translate one to the other, often poorly. I think it would be great for the community if you published both numbers (or at least encouraged it). Clearly this effort has sparked AMD to do a lot more useful shit so I think it VLLM v sglang competition would also probably be good for the community. I would also not allow disable_prefix_cache=True type things but I feel less strongly about that. I also think from VLLM and SGLANG people twitter wars, the fights are not that bad (more publicity) and could actually be resolved more gracefully on this platform.
中文说明
建议为每个模型/硬件配置同时发布 vLLM 和 SGLang 两套基准测试结果,而非仅选择其中一个框架。这有助于社区用户(尤其是使用其中一个框架分支的实验室)更好地对比两者性能,也能促进框架间的良性竞争。同时建议禁止 disable_prefix_cache=True 等不公平配置。
- Dominant language
- Python
- Stars
- 1.7k
- Forks
- 303
- Avg merge
- 1d 13h
- Merged PRs (30d)
- 284
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from SemiAnalysisAI/InferenceX
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
SemiAnalysisAI/InferenceX#2125 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
SemiAnalysisAI/InferenceX#1587 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 78/100
SemiAnalysisAI/InferenceX#1369 · 3 comments ·
-
Difficulty 1/5 1-3 hours Newbie friendliness 76/100
SemiAnalysisAI/InferenceX#1359 · 1 comment ·
-
Difficulty 5/5 Over a week Newbie friendliness 30/100
SemiAnalysisAI/InferenceX#3122 · 3 comments ·
All issues in SemiAnalysisAI/InferenceX
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
bancolombia/sentinel#23 ·
-
test md OpenCI
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
integration:quickjs org:external priority:backlog topic:code-interpreter topic:middleware type:feature
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
langchain-ai/deepagents#6450 ·
-
bug client
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100