intel / intel/llm-scaler

Differences between intel/vllm (latest) and intel/llm-scaler (latest)

Open
#309 8 comments 0 reactions 1 assignee Claimed by @liu-shaojun View on GitHub
vllm
Dominant language
C++
Stars
529
Forks
80
Avg merge
9h 7m
Merged PRs (30d)
38

Description

I tried both and experienced issues with them, although intel/vllm:latest works with gpt-oss-20b.

Summary of my experience:
* Neither worked with Qwen/Qwen3-30B-A3B.
* Neither worked with TP=2 (I have two B60s).
* I’m unsure about the underlying mechanisms (quantisation and memory requirements). Could someone explain them in detail?

Also, Qwen 3.5 series GPTQ-Int4 weights are available. Can llm-scaler support them?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.