cloudnative-pg / cloudnative-pg/plugin-barman-cloud

Sidecar metrics collection over plugin gRPC has no deadline — one stuck operation hangs the whole instance /metrics

Offen
#1,045 3 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
bug
Vorherrschende Sprache
Go
Sterne
191
Forks
72
Ø Merge
2 T. 21 Std.
Gemergte PRs (30 T.)
21

Beschreibung

## Environment

CloudNativePG 1.30.0, plugin-barman-cloud v0.14.0, PostgreSQL 17/18 clusters, VersityGW (posix) S3 endpoint.

## What happens

The instance sidecar collects plugin metrics over the plugin gRPC with **no deadline on the collect path**. Consequence: if any plugin operation gets stuck (in my case a retention/catalog-maintenance run that never completes against a posix-backend S3 gateway — filed separately), the next scrape blocks on the collector and the **entire** `/metrics` endpoint of that instance hangs indefinitely. Prometheus marks the target down and every metric of that instance disappears, not just the plugin's own collectors.

Observed in production: the primary instance's exporter dark for 6.7 hours while the database itself was healthy — WAL-archiver and backup-staleness alerting was blind precisely on the instance where it matters.

## Expected

- A per-collect deadline (a few seconds) on the plugin metrics gRPC call.
- On timeout/error: return the rest of the metrics plus an error counter (e.g. `..._collector_errors_total`), instead of hanging the whole endpoint. A misbehaving plugin should degrade its own metrics, never the instance's.

## Reproduction sketch

1. ObjectStore against VersityGW (posix backend) with any `retentionPolicy` set.
2. Wait for a retention run to start (~30 min cadence) — it never completes against this backend.
3. Scrape the instance sidecar's metrics port: the request hangs until the client gives up; `up` goes to 0 for the instance.

Companion issue describes the retention hang itself; this one is about the missing deadline that turns any such hang into a full metrics outage.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Verfolge den Collect-Pfad der gRPC-Metriken des Plugins im Instance-Sidecar und reproduziere das Problem mit einem feststeckenden Plugin-Vorgang. Untersuche anschließend, wie der /metrics-Endpunkt mit Collector-Fehlern umgeht. Überprüfe, dass ein Timeout pro Collect dem Endpunkt ermöglicht, die verbleibenden Metriken zurückzugeben und einen Fehler zu erfassen, wenn der Plugin-Aufruf hängen bleibt, während funktionierende Plugin-Metriken weiterhin funktionieren.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
go, grpc
Bereich
backend, observability
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Ruhig
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
55/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.