cloudnative-pg / cloudnative-pg/plugin-barman-cloud

Sidecar metrics collection over plugin gRPC has no deadline — one stuck operation hangs the whole instance /metrics

Ouverte
#1,045 3 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
bug
Langage dominant
Go
Étoiles
191
Forks
72
Merge moyen
1 j 16 h
PR mergées (30 j)
18

Description

## Environment

CloudNativePG 1.30.0, plugin-barman-cloud v0.14.0, PostgreSQL 17/18 clusters, VersityGW (posix) S3 endpoint.

## What happens

The instance sidecar collects plugin metrics over the plugin gRPC with **no deadline on the collect path**. Consequence: if any plugin operation gets stuck (in my case a retention/catalog-maintenance run that never completes against a posix-backend S3 gateway — filed separately), the next scrape blocks on the collector and the **entire** `/metrics` endpoint of that instance hangs indefinitely. Prometheus marks the target down and every metric of that instance disappears, not just the plugin's own collectors.

Observed in production: the primary instance's exporter dark for 6.7 hours while the database itself was healthy — WAL-archiver and backup-staleness alerting was blind precisely on the instance where it matters.

## Expected

- A per-collect deadline (a few seconds) on the plugin metrics gRPC call.
- On timeout/error: return the rest of the metrics plus an error counter (e.g. `..._collector_errors_total`), instead of hanging the whole endpoint. A misbehaving plugin should degrade its own metrics, never the instance's.

## Reproduction sketch

1. ObjectStore against VersityGW (posix backend) with any `retentionPolicy` set.
2. Wait for a retention run to start (~30 min cadence) — it never completes against this backend.
3. Scrape the instance sidecar's metrics port: the request hangs until the client gives up; `up` goes to 0 for the instance.

Companion issue describes the retention hang itself; this one is about the missing deadline that turns any such hang into a full metrics outage.

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Suivez le chemin de collect des métriques gRPC du plugin du sidecar d’instance et reproduisez le problème avec une opération du plugin bloquée, puis examinez comment le point de terminaison /metrics gère les erreurs du collector. Vérifiez qu’un délai d’expiration par collect permet au point de terminaison de renvoyer les métriques restantes et d’enregistrer une erreur lorsque l’appel du plugin reste bloqué, tandis que les métriques du plugin en bon état continuent de fonctionner.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
go, grpc
Domaine
backend, observability
Type d'issue
Bug
Difficulté
4/5
Temps estimé
3-5 jours
Activité
Calme
Clarté
Plutôt claire
Accessibilité débutants
55/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.