apache / apache/cloudstack

Connection Pool Leak in Out-of-Band Management (OOB) background task causes Management Server crash (HikariPool MaxActive Exhaustion)

未關閉
#13,382 8 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
component:database component:management-server Severity:Major type:bug
主要語言
Java
星號
3.1k
分支
1.4k
平均合併
7 天 14 小時
30 天內合併 PR
31

描述

### problem

We are experiencing a critical connection leak in Apache CloudStack 4.22.1.0 (running on Ubuntu 24.04). The background task for Out-of-Band Management (OOBM) leaks database connections every time it executes.

Even when setting wait_timeout on MySQL server side or trying to inject pool properties, the connections remain blocked within the Java application state as active, eventually hitting the db.cloud.maxActive threshold (default 250), which causes the management server to stop responding and throw SQLTransientConnectionException.

### versions

CloudStack Version: 4.22.1.0
OS: Ubuntu 24.04.4 LTS
DB: MySQL 8.0.45
Java Version: 17.0.19

### The steps to reproduce the bug

1. Configure Out-of-Band Management (OOBM) for multiple physical hosts
2. Set outofbandmanagement.background.task.execution.interval to a lower value for testing purposes (e.g., 60 or 300 to accelerate the leak
3. Monitor the MySQL SHOW FULL PROCESSLIST; vs. the CloudStack Management Server Metrics over time.
...

### What to do about it?

Every time the OOBM task runs, it opens 1 connection per configured host (3 connections in total for our setup). These connections are never returned to the HikariCP pool (missing .close() or unhandled exception block in the OOBM plugin execution layer).

The Mismatch between DB and Java Pool:

MySQL Side: Sells connections as Sleep. If MySQL kills them via wait_timeout, the sockets are closed on the network layer.

CloudStack/HikariCP Side: Because the leaked connections are still flagged as active (In-Use) by the OOBM thread, HikariCP never runs a health check on them and refuses to evict them via maxLifetime. The internal counter stays at active=250.

Once the counter hits 250, the management server crashes.

Logs & Error Stacktrace:

2026-06-08 21:14:28,070 ERROR [c.c.s.S.ManagementServerCollector] (StatsCollector-1:[ctx-fd90ba07]) (logid:182faa68) Error trying to retrieve management server host statistics com.cloud.utils.exception.CloudRuntimeException: Unable to find on DB, due to: cloud - Connection is not available, request timed out after 30000ms (total=250, active=250, idle=0, waiting=22)

Is there any workaround available? Maybe switching to dbcp?

貢獻指南

開啟貢獻指南

研究方向

先追蹤 OOBM 背景工作執行,以及其資料庫連線取得與釋放路徑;issue 未指定特定檔案或測試。透過縮短執行間隔進行重現,同時比較 MySQL 程序清單項目與 CloudStack pool 指標。完成的判定標準是:工作重複執行後會將連線歸還 pool,且作用中連線數不再持續增加至設定的最大值。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
java, mysql
領域
backend, database, infrastructure
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
活躍
描述清晰度
基本清楚
新手友好度
48/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。