Performance Regression or Improvement: test_cloudml_benchmark_cirteo_no_shuffle_10GB-runtime_sec:runtime_sec
- Dominant language
- Java
- Stars
- 8.7k
- Forks
- 4.7k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 196
Description
Performance change found in the
test: `test_cloudml_benchmark_cirteo_no_shuffle_10GB-runtime_sec` for the metric: `runtime_sec`.
For more information on how to triage the alerts, please look at
`Triage performance alert issues` section of the [README](https://github.com/apache/beam/tree/master/sdks/python/apache_beam/testing/analyzers/README.md#triage-performance-alert-issues).
`Test description:` TFT Criteo test on 10 GB data with no Reshuffle.
Test link - [Test link](https://github.com/apache/beam/blob/42d0a6e3564d8b9c5d912428a6de18fb22a13ac1/sdks/python/apache_beam/testing/benchmarks/cloudml/cloudml_benchmark_test.py#L82)
```
timestamp: Mon Sep 7 23:19:16 2026, metric_value: 2998.61
timestamp: Sun Sep 6 23:09:38 2026, metric_value: 3075.77
timestamp: Sat Sep 5 23:39:23 2026, metric_value: 3069.05
timestamp: Fri Sep 4 23:08:06 2026, metric_value: 2910.89
timestamp: Thu Sep 3 23:05:58 2026, metric_value: 3123.24
timestamp: Wed Sep 2 23:00:42 2026, metric_value: 3080.73 <---- Anomaly
timestamp: Tue Sep 1 23:23:12 2026, metric_value: 2366.35
timestamp: Mon Aug 31 22:52:36 2026, metric_value: 2537.10
timestamp: Mon Aug 31 20:43:04 2026, metric_value: 2407.92
timestamp: Sun Aug 30 22:39:01 2026, metric_value: 2299.40
timestamp: Sat Aug 29 23:00:03 2026, metric_value: 2421.35
timestamp: Fri Aug 28 23:58:09 2026, metric_value: 2360.65
timestamp: Fri Aug 28 01:54:40 2026, metric_value: 2298.42
timestamp: Wed Aug 26 23:28:17 2026, metric_value: 2362.46
timestamp: Wed Aug 26 12:11:02 2026, metric_value: 2356.62
timestamp: Wed Aug 19 22:04:50 2026, metric_value: 2308.39
```
Contributor guide
Research direction
Read the Triage performance alert issues section in sdks/python/apache_beam/testing/analyzers/README.md, then inspect the test at sdks/python/apache_beam/testing/benchmarks/cloudml/cloudml_benchmark_test.py:82. Compare the reported runtime anomaly with the surrounding measurements and investigate the cause. Done means the regression is explained and the benchmark result is validated.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100