ByteDance-Seed / ByteDance-Seed/Triton-distributed
Performance issue about A2A on 8xH800
Open
- Dominant language
- Python
- Stars
- 1.5k
- Forks
- 172
- PR merge metrics
- No merged PRs in 30d
Description
as the seq length grow,the performance of low_latency_all_to_all kernel provided is worse than that in pytorch
with this config
Contributor guide
Research direction
Reproduce the reported A2A benchmark on 8xH800 using the configuration shown, comparing the low_latency_all_to_all kernel with PyTorch as sequence length increases. Trace the relevant benchmark and kernel entry points, then establish the cause of the performance crossover and verify that the affected case no longer underperforms PyTorch.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100