prometheus / prometheus/client_java

OpenTelemetry bridge: classic histograms exported with +Inf boundary included, violating counts == boundaries + 1 contract

Abierto Apto para principiantes
#2,416 1 comentario 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Lenguaje dominante
Java
Estrellas
2.3k
Forks
833
Merge medio
2 d 16 h
PR fusionados (30 d)
86

Descripción

Summary

PrometheusClassicHistogram.toOtelDataPoint() converts a Prometheus classic histogram into an OpenTelemetry HistogramPointData, but it passes the Prometheus bucket boundaries including the final +Inf boundary and emits exactly one count per boundary. The OpenTelemetry data model requires counts.size() == boundaries.size() + 1, with boundaries containing only finite values (the last, +Inf bucket is implicit). The converted point therefore violates the OTel contract that downstream consumers rely on.

Static-analysis finding against current master; not executed here.

Location

  • File: prometheus-metrics-exporter-opentelemetry/src/main/java/io/prometheus/metrics/exporter/opentelemetry/otelmodel/PrometheusClassicHistogram.java
  • Methods: makeBoundaries(ClassicHistogramBuckets) (~lines 66-72) and makeCounts(ClassicHistogramBuckets) (~lines 74-80), called from toOtelDataPoint() (~lines 45-58).

Problem

makeBoundaries() iterates all classic buckets and appends every upper bound - for a Prometheus classic histogram the last bucket's upper bound is Double.POSITIVE_INFINITY, so +Inf ends up inside HistogramPointData.getBoundaries(). makeCounts() returns buckets.size() counts, i.e. counts.size() == boundaries.size().

Per the OTel SDK data model (io.opentelemetry.sdk.metrics.data.HistogramPointData), boundaries are the finite bucket edges and there is always one extra count for the implicit last bucket:

  • getCounts().size() == getBoundaries().size() + 1
  • boundaries must be finite (Double.POSITIVE_INFINITY is not a legal explicit bound; the OTLP explicit_bounds field likewise expects finite values)

Concretely, a Prometheus histogram with bounds [1, 5, +Inf] and counts [c0, c1, c2] is converted to:

  • boundaries [1.0, 5.0, Infinity]
  • counts [c0, c1, c2]

An OTel consumer reading this point interprets it as four buckets: [..,1], [1,5], [5,+Inf] (count c2) plus an implicit trailing (+Inf) bucket with count 0 - misrepresenting both bucket structure and totals. Anything that validates the model (or the OTLP exporter's serialization of infinite explicit bounds) will reject or distort the histogram instead of exporting it faithfully.

Trigger / Reproduction

Based on static analysis; no runtime run was performed:

  • Expose any classic histogram through PrometheusMetricsExporter with the OpenTelemetry bridge enabled (MetricDataFactory line ~73 constructs PrometheusClassicHistogram whenever the snapshot has classic histogram data).
  • Observe the resulting HistogramPointData: getBoundaries().contains(Double.POSITIVE_INFINITY) and getCounts().size() == getBoundaries().size().

Expected Behavior

The conversion should drop the final +Inf upper bound from boundaries and append the total observation count as the extra last element of counts (the implicit overflow bucket), e.g. boundaries [1.0, 5.0], counts [c0, c1, c0+c1+c2]. (calculateCount() already computes this sum when the snapshot lacks an explicit count.)

Actual Behavior

+Inf remains in the boundary list and no implicit-bucket count is appended, producing an out-of-contract HistogramPointData.

Impact

Every classic histogram exported through the OpenTelemetry bridge carries malformed bucket metadata: OTLP exports can fail validation or silently shift all bucket assignments, and aggregation logic built on the OTel model (e.g., heatmap rendering, downstream collectors) computes wrong distributions.

Suggested Direction

In toOtelDataPoint() (or in the two helpers), strip the trailing +Inf boundary when present and build counts as the per-boundary counts followed by the total count. Keeping the helpers symmetric (boundaries.size() == counts.size() - 1) would make the invariant locally verifiable.

Evidence

  • PrometheusClassicHistogram.makeBoundaries(): unconditional buckets.getUpperBound(i) loop; ClassicHistogramBuckets snapshots from the core module always end with +Inf for classic histograms.
  • MetricDataFactory.java line ~73: result is fed directly into the OTel MetricData tree consumed by the exporter pipeline.

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Empieza en prometheus-metrics-exporter-opentelemetry/src/main/java/io/prometheus/metrics/exporter/opentelemetry/otelmodel/PrometheusClassicHistogram.java, leyendo toOtelDataPoint(), makeBoundaries() y makeCounts(); después sigue MetricDataFactory.java para entender el punto de entrada. Verifica que la HistogramPointData resultante tenga únicamente límites finitos y que counts.size() == boundaries.size() + 1, donde el recuento final representa el bucket de desbordamiento implícito.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
java
Área
observability
Tipo de issue
Error
Dificultad
2/5
Tiempo estimado
1-3 horas
Estado de actividad
Activo
Claridad
Bien especificado
Aptitud para principiantes
82/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.