alibaba / alibaba/BladeDISC

[TorchBench] Performance Signal Detected

Open
#1,237 0 comments 0 reactions 0 assignees View on GitHub
Benchmark
Dominant language
C++
Stars
933
Forks
169
PR merge metrics
No merged PRs in 30d

Description

TorchBench CI has detected a performance signal.

Affected Tests:

- eval-cuda-fp32:
- hf_Bert[disc (latency)] 8.16 -> 13.834, -69.5343%
- hf_Bert[dynamo-disc (latency)] 6.865 -> 6.175, +10.051%
- hf_Bert[disc (compiled)] 1151 -> 0
- hf_Bert[disc (clusters)] 1 -> 0
- eval-cuda-fp16:
- hf_Bert[disc (compiled)] 1151 -> 0
- hf_Bert[disc (clusters)] 1 -> 0

detail data can be seen in oss://bladedisc-ci/TorchBench/gpu/tiny/20230803-15
created by TorchBench CI automatically

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the TorchBench CI report and compare the affected hf_Bert cases in eval-cuda-fp32 and eval-cuda-fp16 against the detail data at oss://bladedisc-ci/TorchBench/gpu/tiny/20230803-15. Determine whether the latency, compiled, and clusters changes are reproducible; done means the signal is explained and the necessary follow-up is recorded.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, pytorch
Domain
compilers, machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.