apache / apache/cloudstack

prometheus improvements

Open
#13,667 4 comments 0 reactions 0 assignees View on GitHub
type:technical-debt
Dominant language
Java
Stars
3.1k
Forks
1.4k
Avg merge
6d 19h
Merged PRs (30d)
32

Description

### problem

1. Give the exporter's HttpServer an explicit bounded executor (httpServer.setExecutor(Executors.newFixedThreadPool(2))) so one slow scrape can't serialize/queue all others.
2. Add a short TTL/in-flight guard around updateMetrics() (e.g., skip recompute if last run was < N seconds ago, or synchronize so concurrent scrapes share one in-progress computation) so scrape frequency can never multiply backend load.
3. Instrument: log/measure updateMetrics() wall-clock time so the reporter (and CI) can confirm which sub-metric collector is actually slow and verify the fix closes the growth.

additional comments:

1. Stale dynamic config (CONFIRMED) — capacity.calculate.workers is a runtime-dynamic setting, but the new shared executor only reads it once at first creation; live changes are silently ignored until a restart.
2. Swallowed capacity-recalculation abort (CONFIRMED per the extra verify pass) — shutdown racing an in-flight recalculation throws RejectedExecutionException, caught by the blanket catch(Throwable) in recalculateCapacity(), silently skipping storage/IP/VLAN updates for that cycle.
3. Unsynchronized race on _capacityExecutorService (CONFIRMED per the extra verify pass) — can leak a freshly-recreated pool that's never shut down again.
4. Shared fixed-size pool serializes previously-independent callers (PLAUSIBLE) — rolling-maintenance host-drain gating can now queue behind the hourly timer or API-triggered recalculations.
5. Pool no longer bounded to actual task count, so it can stay oversized/stale relative to fleet size (efficiency).
6. Bundling this executor-lifecycle rewrite into what the reported bug (#13586) only needed a one-line fix for (altitude/scope creep).
7. Inconsistent lazy-vs-eager thread-pool lifecycle pattern within the same class (reuse/convention).
8. Minor: the synchronized getter is called per-loop-iteration instead of hoisted once (efficiency).

Contributor guide

Open the contributing guide

Research direction

Locate the exporter HttpServer, updateMetrics(), recalculateCapacity(), and _capacityExecutorService; read their executor lifecycle and capacity.calculate.workers handling first. Trace concurrent scrapes, shutdown races, and rolling-maintenance, API, and timer callers. Done means the concurrency, dynamic configuration, abort handling, and timing instrumentation are explicit and verifiable without silently losing recalculations.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, prometheus
Domain
infrastructure, observability
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.