prometheus / prometheus/client_java

OpenTelemetry bridge: classic histograms exported with +Inf boundary included, violating counts == boundaries + 1 contract

Offen Anfängerfreundlich
#2,416 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Vorherrschende Sprache
Java
Sterne
2.3k
Forks
833
Ø Merge
2 T. 16 Std.
Gemergte PRs (30 T.)
86

Beschreibung

Summary

PrometheusClassicHistogram.toOtelDataPoint() converts a Prometheus classic histogram into an OpenTelemetry HistogramPointData, but it passes the Prometheus bucket boundaries including the final +Inf boundary and emits exactly one count per boundary. The OpenTelemetry data model requires counts.size() == boundaries.size() + 1, with boundaries containing only finite values (the last, +Inf bucket is implicit). The converted point therefore violates the OTel contract that downstream consumers rely on.

Static-analysis finding against current master; not executed here.

Location

  • File: prometheus-metrics-exporter-opentelemetry/src/main/java/io/prometheus/metrics/exporter/opentelemetry/otelmodel/PrometheusClassicHistogram.java
  • Methods: makeBoundaries(ClassicHistogramBuckets) (~lines 66-72) and makeCounts(ClassicHistogramBuckets) (~lines 74-80), called from toOtelDataPoint() (~lines 45-58).

Problem

makeBoundaries() iterates all classic buckets and appends every upper bound - for a Prometheus classic histogram the last bucket's upper bound is Double.POSITIVE_INFINITY, so +Inf ends up inside HistogramPointData.getBoundaries(). makeCounts() returns buckets.size() counts, i.e. counts.size() == boundaries.size().

Per the OTel SDK data model (io.opentelemetry.sdk.metrics.data.HistogramPointData), boundaries are the finite bucket edges and there is always one extra count for the implicit last bucket:

  • getCounts().size() == getBoundaries().size() + 1
  • boundaries must be finite (Double.POSITIVE_INFINITY is not a legal explicit bound; the OTLP explicit_bounds field likewise expects finite values)

Concretely, a Prometheus histogram with bounds [1, 5, +Inf] and counts [c0, c1, c2] is converted to:

  • boundaries [1.0, 5.0, Infinity]
  • counts [c0, c1, c2]

An OTel consumer reading this point interprets it as four buckets: [..,1], [1,5], [5,+Inf] (count c2) plus an implicit trailing (+Inf) bucket with count 0 - misrepresenting both bucket structure and totals. Anything that validates the model (or the OTLP exporter's serialization of infinite explicit bounds) will reject or distort the histogram instead of exporting it faithfully.

Trigger / Reproduction

Based on static analysis; no runtime run was performed:

  • Expose any classic histogram through PrometheusMetricsExporter with the OpenTelemetry bridge enabled (MetricDataFactory line ~73 constructs PrometheusClassicHistogram whenever the snapshot has classic histogram data).
  • Observe the resulting HistogramPointData: getBoundaries().contains(Double.POSITIVE_INFINITY) and getCounts().size() == getBoundaries().size().

Expected Behavior

The conversion should drop the final +Inf upper bound from boundaries and append the total observation count as the extra last element of counts (the implicit overflow bucket), e.g. boundaries [1.0, 5.0], counts [c0, c1, c0+c1+c2]. (calculateCount() already computes this sum when the snapshot lacks an explicit count.)

Actual Behavior

+Inf remains in the boundary list and no implicit-bucket count is appended, producing an out-of-contract HistogramPointData.

Impact

Every classic histogram exported through the OpenTelemetry bridge carries malformed bucket metadata: OTLP exports can fail validation or silently shift all bucket assignments, and aggregation logic built on the OTel model (e.g., heatmap rendering, downstream collectors) computes wrong distributions.

Suggested Direction

In toOtelDataPoint() (or in the two helpers), strip the trailing +Inf boundary when present and build counts as the per-boundary counts followed by the total count. Keeping the helpers symmetric (boundaries.size() == counts.size() - 1) would make the invariant locally verifiable.

Evidence

  • PrometheusClassicHistogram.makeBoundaries(): unconditional buckets.getUpperBound(i) loop; ClassicHistogramBuckets snapshots from the core module always end with +Inf for classic histograms.
  • MetricDataFactory.java line ~73: result is fed directly into the OTel MetricData tree consumed by the exporter pipeline.

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne in prometheus-metrics-exporter-opentelemetry/src/main/java/io/prometheus/metrics/exporter/opentelemetry/otelmodel/PrometheusClassicHistogram.java, lies toOtelDataPoint(), makeBoundaries() und makeCounts(); verfolge dann MetricDataFactory.java, um den Einstiegspunkt zu verstehen. Überprüfe, dass die resultierende HistogramPointData nur endliche Grenzen enthält und counts.size() == boundaries.size() + 1 gilt, wobei der letzte Zähler den impliziten Überlauf-Bucket repräsentiert.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
java
Bereich
observability
Issue-Typ
Bug
Schwierigkeit
2/5
Geschätzter Aufwand
1-3 Stunden
Aktivitätsstatus
Aktiv
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
82/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.