apache / apache/cloudstack

prometheus improvements

未关闭
#13,667 4 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
type:technical-debt
主要语言
Java
星标
3.1k
派生
1.4k
平均合并
6 天 19 小时
30 天内合并 PR
32

描述

### problem

1. Give the exporter's HttpServer an explicit bounded executor (httpServer.setExecutor(Executors.newFixedThreadPool(2))) so one slow scrape can't serialize/queue all others.
2. Add a short TTL/in-flight guard around updateMetrics() (e.g., skip recompute if last run was < N seconds ago, or synchronize so concurrent scrapes share one in-progress computation) so scrape frequency can never multiply backend load.
3. Instrument: log/measure updateMetrics() wall-clock time so the reporter (and CI) can confirm which sub-metric collector is actually slow and verify the fix closes the growth.

additional comments:

1. Stale dynamic config (CONFIRMED) — capacity.calculate.workers is a runtime-dynamic setting, but the new shared executor only reads it once at first creation; live changes are silently ignored until a restart.
2. Swallowed capacity-recalculation abort (CONFIRMED per the extra verify pass) — shutdown racing an in-flight recalculation throws RejectedExecutionException, caught by the blanket catch(Throwable) in recalculateCapacity(), silently skipping storage/IP/VLAN updates for that cycle.
3. Unsynchronized race on _capacityExecutorService (CONFIRMED per the extra verify pass) — can leak a freshly-recreated pool that's never shut down again.
4. Shared fixed-size pool serializes previously-independent callers (PLAUSIBLE) — rolling-maintenance host-drain gating can now queue behind the hourly timer or API-triggered recalculations.
5. Pool no longer bounded to actual task count, so it can stay oversized/stale relative to fleet size (efficiency).
6. Bundling this executor-lifecycle rewrite into what the reported bug (#13586) only needed a one-line fix for (altitude/scope creep).
7. Inconsistent lazy-vs-eager thread-pool lifecycle pattern within the same class (reuse/convention).
8. Minor: the synchronized getter is called per-loop-iteration instead of hoisted once (efficiency).

贡献指南

打开贡献指南

调研方向

定位 exporter HttpServer、updateMetrics()、recalculateCapacity() 和 _capacityExecutorService;先阅读它们的执行器生命周期以及 capacity.calculate.workers 的处理方式。跟踪并发 scrape、关闭竞争,以及 rolling-maintenance、API 和计时器调用方。完成意味着并发性、动态配置、abort 处理和时间测量 instrumentation 都是明确且可验证的,同时不会静默丢失重新计算。

由索引模型根据 Issue 内容生成。

评估

技术栈
java, prometheus
领域
infrastructure, observability
Issue 类型
重构
难度
5/5
预计耗时
一周以上
活跃度
冷清
描述清晰度
需要澄清
新手友好度
28/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。