ByteDance-Seed / ByteDance-Seed/Triton-distributed

Performance issue about A2A on 8xH800

Open
#98 8 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.5k
Forks
172
PR merge metrics
No merged PRs in 30d

Description

as the seq length grow,the performance of low_latency_all_to_all kernel provided is worse than that in pytorch
Image
with this config

Image

Contributor guide

Open the contributing guide

Research direction

Reproduce the reported A2A benchmark on 8xH800 using the configuration shown, comparing the low_latency_all_to_all kernel with PyTorch as sequence length increases. Trace the relevant benchmark and kernel entry points, then establish the cause of the performance crossover and verify that the affected case no longer underperforms PyTorch.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.