sgl-project / sgl-project/SpecForge

[Question]What factors contribute to Eagle3's reduced acceleration performance on VLM architectures compared to traditional LLMs?

Open
#352 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
1.2k
Forks
346
Avg merge
4d 1h
Merged PRs (30d)
41

Description

First, sincere thanks to the SGLang community for enabling the rapid deployment of Eagle3 on state-of-the-art VLMs such as Qwen2.5-VL.

In my preliminary tests conducted under identical settings, however, the speed-up brought by Eagle3 is noticeably smaller than that observed on language models (more detailed plz see https://github.com/sgl-project/sglang/pull/8801#issuecomment-3615087060).

Could you kindly share any insights into what might be causing this gap? Any guidance would be highly appreciated :)

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the linked SGLang pull-request comment and compare the reported Eagle3 results for Qwen2.5-VL with the identical-settings language-model results. The issue is done when the factors behind the smaller VLM speed-up are identified and documented clearly enough to guide further investigation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
ai, machine-learning, performance
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.