intel / intel/llm-scaler

Clarify vllm commitment plan that was announced

Open
#336 10 comments 5 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
529
Forks
80
Avg merge
9h 7m
Merged PRs (30d)
38

Description

Given this timeline of announcements:

- Nov 11 2025: [link](https://vllm.ai/blog/intel-arc-pro-b) **Fast and Affordable LLMs serving on Intel Arc Pro B-Series GPUs with vLLM**
- Mar 11 2026: [link](https://github.com/intel/llm-scaler/releases/tag/vllm-0.14.0-b8.1) **[2026.03] We released intel/llm-scaler-vllm:0.14.0-b8.1 to support Qwen3.5-27B, Qwen3.5-35B-A3B and Qwen3.5-122B-A10B (FP8/INT4 online quantization, GPTQ)**
- Mar 25 2026: [link](https://newsroom.intel.com/client-computing/intel-core-ultra-series-3-with-vpro-powers-next-gen-pcs-on-18a) **Intel Arc Pro B70 and B65 discrete GPUs expand Intel’s professional graphics portfolio**

And the statement of commitment:
> We commit to deepening the integration between our optimizations and the core vLLM project. Our roadmap includes providing full support for upstream vLLM features, delivering state-of-the-art performance optimizations for a broad range of models, with a special focus on popular, high-performance LLMs on Intel® hardware, and actively contributing our enhancements back to the vLLM upstream community.

Yet Intel released the latest support for Qwen 3.5 via the llm-scaler fork of vllm, using a **_2 month old version of vllm._**

**_Please explain why this happened and what is the plan to meet your stated commitment?_**

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.