dmlc / dmlc/xgboost

[CI] Analysis of recent increase in CI costs

Open
#12,152 3 comments 0 reactions 1 assignee Claimed by @chyunsu3 View on GitHub
Dominant language
C++
Stars
28.8k
Forks
8.9k
Avg merge
1d 12h
Merged PRs (30d)
54

Description

Starting from 3 months ago, the CI cost skyrocketed.

Image

(Each color bar indicates the workload type, as defined in [here](https://github.com/dmlc/xgboost/blob/master/.github/runs-on.yml). Each workload maps to one or more pre-defined EC2 instance types. Here, we use workload classification for better legibility.)

I've ran through some numbers from the AWS Cost Explorer. Here is the finding:

**We are consuming more CI minutes, and each CI minute costs more.**

Both the run time and the unit cost have gone up simultaneously.

Image

**Why has the unit cost gone up?**

The rise in the unit cost mostly coincides with the move away from spot instances. More specifically,
* The unit cost of GPU workloads went up on August 2025, when GPU workloads began to use on-demand instances instead of spots (#11644)
* The unit cost of CPU workloads went up on November 2025, when CPU workloads began to use on-demand instances instead of spots (#11824)

The CI pipelines used to utilize [spot instances](https://aws.amazon.com/ec2/spot/), which uses slack (idle) capacity of AWS EC2 in order to secure discounts. We decided to stop using spots because they tend to get interrupted often. See discussion in #11472.

While [AWS raised prices for certain high-end GPU machines](https://www.datacenterdynamics.com/en/news/aws-quietly-increases-prices-for-h200-ec2-instances-by-15/), I cannot locate any evidence that AWS significantly raised prices for the low-end machines we use for the CI. So we should attribute most of the cost increase to the adoption of on-demand instances.

**What can we do to reduce the cost?**

Looking from the plots, I identify **two major bottlenecks**:
* `linux-amd64-mgpu`: The cost problem is the worst with this workload, as we have multiple confounding issues. As a result, the workload now costs **3-4x** compared to the 2025 baseline level.
- The workload uses the most costly EC2 instance (`g4dn.12xlarge`), which costs whopping 3.9 USD per hour.
- The unit price doubled since the adoption of on-demand pricing.
- The workload now consumes **2-2.5x minutes** than the baseline level from the 2025 baseline.
* `linux-amd64-cpu`: This workload consumes 1.5-2x more minutes from the 2025 baseline level. The unit price went up by 50%, culminating to the total cost increase of **2-2.5x**.

Other workloads (e.g. Windows, ARM) take comparatively insignificant portions of the compute budget.

For the two bottlenecks above, we should adopt the following remedies:

* Require manual approvals for these two workloads to run. This will also protect us from sudden cost increases owing to vibe-coded spam pull requests (*).
* Slim down test suites, to make them run faster.

(*) This should be distinguished from good quality pull requests where the contributor judiciously uses the LLM agent and properly vets its output for legibility and correctness.

## Appendix

**Cost (USD) by workload type**

  | linux-amd64-cpu | linux-amd64-gpu | linux-amd64-mgpu | linux-arm64-cpu | linux-arm64-gpu | windows-amd64-cpu | windows-gpu
-- | -- | -- | -- | -- | -- | -- | --
2025/04 | 174.2629 | 10.3491 | 139.4857 | 9.393105 | 0 | 122.0691 | 28.68151
2025/05 | 205.6721 | 12.19294 | 156.9493 | 9.469875 | 0 | 124.649 | 28.43021
2025/06 | 212.6446 | 10.63133 | 274.4651 | 12.6765 | 0 | 105.4018 | 26.74545
2025/07 | 357.6044 | 22.10261 | 285.0447 | 20.49686 | 0 | 207.2675 | 46.18852
2025/08 | 293.7534 | 28.51596 | 289.9386 | 18.04933 | 0 | 130.8959 | 50.67155
2025/09 | 225.8608 | 49.0273 | 331.417 | 13.02594 | 0 | 102.4939 | 52.80956
2025/10 | 301.4604 | 63.7274 | 395.3641 | 16.23963 | 0 | 139.2879 | 54.19338
2025/11 | 322.8928 | 65.19376 | 390.0916 | 20.34282 | 0 | 167.1499 | 57.63707
2025/12 | 440.1882 | 75.95425 | 428.6433 | 62.20534 | 43.62937 | 205.0661 | 62.15937
2026/01 | 838.7721 | 174.0611 | 1026.89 | 136.9029 | 122.9327 | 447.002 | 142.6108
2026/02 | 538.4042 | 137.8891 | 884.7216 | 78.22418 | 96.13124 | 331.783 | 119.1512
2026/03 | 713.1427 | 185.4974 | 1155.252 | 110.9274 | 133.6279 | 394.1583 | 162.3393

**Usage (Hours) by workload type**

  | linux-amd64-cpu | linux-amd64-gpu | linux-amd64-mgpu | linux-arm64-cpu | linux-arm64-gpu | windows-amd64-cpu | windows-gpu
-- | -- | -- | -- | -- | -- | -- | --
2025/04 | 466.9384 | 42.79666 | 76.21918 | 37.76695 | 0 | 61.68499 | 37.43306
2025/05 | 524.4271 | 52.44445 | 86.76195 | 44.69862 | 0 | 64.9 | 40.36028
2025/06 | 480.1265 | 41.46279 | 107.5695 | 48.86557 | 0 | 56.32724 | 37.28361
2025/07 | 900.1877 | 82.57502 | 145.2772 | 80.36 | 0 | 119.8503 | 69.54667
2025/08 | 751.617 | 78.58666 | 103.1603 | 64.86612 | 0 | 69.26278 | 57.65861
2025/09 | 600.9434 | 93.20779 | 84.71805 | 43.89363 | 0 | 48.92167 | 47.15139
2025/10 | 739.8265 | 121.1547 | 101.0644 | 47.60861 | 0 | 65.84528 | 48.38695
2025/11 | 694.5032 | 123.9425 | 99.71666 | 55.74 | 0 | 73.9525 | 51.46167
2025/12 | 714.5913 | 144.3997 | 109.5714 | 114.3481 | 103.8794 | 75.83805 | 55.49944
2026/01 | 1361.643 | 329.7219 | 262.4975 | 251.6597 | 292.6969 | 165.3114 | 127.3311
2026/02 | 874.0328 | 262.1467 | 226.1558 | 143.7944 | 228.8839 | 122.7008 | 106.385
2026/03 | 1157.699 | 352.6567 | 295.3097 | 203.9106 | 318.1617 | 145.7686 | 144.9458

**Unit cost (USD/hour) by workload type**

  | linux-amd64-cpu | linux-amd64-gpu | linux-amd64-mgpu | linux-arm64-cpu | linux-arm64-gpu | windows-amd64-cpu | windows-gpu
-- | -- | -- | -- | -- | -- | -- | --
2025/04 | 0.373203 | 0.24182 | 1.830061 | 0.248712 | N/A | 1.978912 | 0.766208
2025/05 | 0.392184 | 0.232492 | 1.808964 | 0.211861 | N/A | 1.920631 | 0.704411
2025/06 | 0.442893 | 0.256406 | 2.551516 | 0.259416 | N/A | 1.87124 | 0.717351
2025/07 | 0.397255 | 0.267667 | 1.962074 | 0.255063 | N/A | 1.729387 | 0.664137
2025/08 | 0.390829 | 0.36286 | 2.810565 | 0.278255 | N/A | 1.889845 | 0.87882
2025/09 | 0.375844 | 0.526 | 3.912 | 0.296761 | N/A | 2.095061 | 1.12
2025/10 | 0.407474 | 0.526 | 3.912 | 0.341107 | N/A | 2.115382 | 1.12
2025/11 | 0.464926 | 0.526 | 3.912 | 0.364959 | N/A | 2.260233 | 1.12
2025/12 | 0.616 | 0.526 | 3.912 | 0.544 | 0.42 | 2.704 | 1.12
2026/01 | 0.616 | 0.527903 | 3.912 | 0.544 | 0.42 | 2.704 | 1.12
2026/02 | 0.616 | 0.526 | 3.912 | 0.544 | 0.42 | 2.704 | 1.12
2026/03 | 0.616 | 0.526 | 3.912 | 0.544 | 0.42 | 2.704 | 1.12

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.