ByteDance-Seed / ByteDance-Seed/Triton-distributed

Any comparison between mega-triton-kernel & sglang/vllm?

Open
#100 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.5k
Forks
172
PR merge metrics
No merged PRs in 30d

Description

https://github.com/ByteDance-Seed/Triton-distributed/blob/main/docs/mega_triton_kernel.md

Does this mega triton kernel implementation provide critical inference engine caps, say:
- paged attention
- continuous batching
- zero scheduling overheads

What's the goal of mega-triton-kernel? Is it aimed to replace the existing vllm/sglang op+graph based impl?

Contributor guide

Open the contributing guide

Research direction

Start with docs/mega_triton_kernel.md and review the documented scope of mega-triton-kernel. Compare its stated capabilities with paged attention, continuous batching, and scheduling overhead, then document whether those capabilities are provided and clarify its relationship to vLLM and SGLang.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.