apache / apache/beam

Performance Regression or Improvement: test_cloudml_benchmark_cirteo_no_shuffle_10GB-runtime_sec:runtime_sec

Open
#40,067 1 comment 0 reactions 0 assignees View on GitHub
awaiting triage perf-alert
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

Performance change found in the
test: `test_cloudml_benchmark_cirteo_no_shuffle_10GB-runtime_sec` for the metric: `runtime_sec`.

For more information on how to triage the alerts, please look at
`Triage performance alert issues` section of the [README](https://github.com/apache/beam/tree/master/sdks/python/apache_beam/testing/analyzers/README.md#triage-performance-alert-issues).

`Test description:` TFT Criteo test on 10 GB data with no Reshuffle.
Test link - [Test link](https://github.com/apache/beam/blob/42d0a6e3564d8b9c5d912428a6de18fb22a13ac1/sdks/python/apache_beam/testing/benchmarks/cloudml/cloudml_benchmark_test.py#L82)

```

timestamp: Mon Sep 7 23:19:16 2026, metric_value: 2998.61
timestamp: Sun Sep 6 23:09:38 2026, metric_value: 3075.77
timestamp: Sat Sep 5 23:39:23 2026, metric_value: 3069.05
timestamp: Fri Sep 4 23:08:06 2026, metric_value: 2910.89
timestamp: Thu Sep 3 23:05:58 2026, metric_value: 3123.24
timestamp: Wed Sep 2 23:00:42 2026, metric_value: 3080.73 <---- Anomaly
timestamp: Tue Sep 1 23:23:12 2026, metric_value: 2366.35
timestamp: Mon Aug 31 22:52:36 2026, metric_value: 2537.10
timestamp: Mon Aug 31 20:43:04 2026, metric_value: 2407.92
timestamp: Sun Aug 30 22:39:01 2026, metric_value: 2299.40
timestamp: Sat Aug 29 23:00:03 2026, metric_value: 2421.35
timestamp: Fri Aug 28 23:58:09 2026, metric_value: 2360.65
timestamp: Fri Aug 28 01:54:40 2026, metric_value: 2298.42
timestamp: Wed Aug 26 23:28:17 2026, metric_value: 2362.46
timestamp: Wed Aug 26 12:11:02 2026, metric_value: 2356.62
timestamp: Wed Aug 19 22:04:50 2026, metric_value: 2308.39

```

Contributor guide

Open the contributing guide

Research direction

Read the Triage performance alert issues section in sdks/python/apache_beam/testing/analyzers/README.md, then inspect the test at sdks/python/apache_beam/testing/benchmarks/cloudml/cloudml_benchmark_test.py:82. Compare the reported runtime anomaly with the surrounding measurements and investigate the cause. Done means the regression is explained and the benchmark result is validated.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.