Differences between intel/vllm (latest) and intel/llm-scaler (latest)
Open
vllm
- Dominant language
- C++
- Stars
- 529
- Forks
- 80
- Avg merge
- 9h 7m
- Merged PRs (30d)
- 38
Description
I tried both and experienced issues with them, although intel/vllm:latest works with gpt-oss-20b.
Summary of my experience:
* Neither worked with Qwen/Qwen3-30B-A3B.
* Neither worked with TP=2 (I have two B60s).
* I’m unsure about the underlying mechanisms (quantisation and memory requirements). Could someone explain them in detail?
Also, Qwen 3.5 series GPTQ-Int4 weights are available. Can llm-scaler support them?
Contributor guide
Assessment
This issue has not been assessed yet.