ByteDance-Seed / ByteDance-Seed/Triton-distributed
Any comparison between mega-triton-kernel & sglang/vllm?
- Dominant language
- Python
- Stars
- 1.5k
- Forks
- 172
- PR merge metrics
- No merged PRs in 30d
Description
https://github.com/ByteDance-Seed/Triton-distributed/blob/main/docs/mega_triton_kernel.md
Does this mega triton kernel implementation provide critical inference engine caps, say:
- paged attention
- continuous batching
- zero scheduling overheads
What's the goal of mega-triton-kernel? Is it aimed to replace the existing vllm/sglang op+graph based impl?
Contributor guide
Research direction
Start with docs/mega_triton_kernel.md and review the documented scope of mega-triton-kernel. Compare its stated capabilities with paged attention, continuous batching, and scheduling overhead, then document whether those capabilities are provided and clarify its relationship to vLLM and SGLang.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100