cloudnative-pg / cloudnative-pg/plugin-barman-cloud
Retention run never completes against a POSIX-backend S3 gateway (VersityGW) and stalls the sidecar metrics endpoint
- Langage dominant
- Go
- Étoiles
- 191
- Forks
- 72
- Merge moyen
- 1 j 16 h
- PR mergées (30 j)
- 18
Description
## Environment
- CloudNativePG 1.30.0, plugin-barman-cloud v0.14.0 (sidecar image `plugin-barman-cloud-sidecar:v0.14.0`)
- PostgreSQL 18.4 and 17.10 clusters (3 instances each)
- Object store: [VersityGW](https://github.com/versity/versitygw) S3 gateway with the **posix** backend, plain HTTP endpoint inside the cluster
- Base backups and WAL archiving against this endpoint work fine (nightly base from standby, continuous WAL, PITR restore drills pass)
## What happens
With any `retentionPolicy` set on the `ObjectStore` (I used `"30d"`), the catalog-maintenance/retention run starts on its ~30 minute cadence and **never completes**. There is no error surfaced anywhere — no failed condition on the ObjectStore, nothing actionable in the sidecar logs — the run just stays stuck, and a new one piles up on the next cadence.
The visible damage is on the metrics side: while a retention run is stuck, the instance sidecar's `/metrics` endpoint stops responding (scrapes time out; see companion issue about the missing deadline on the metrics path). In practice the exporter for the affected instance goes dark for hours — in my case the *primary* was unscrapeable for 6.7 hours while PostgreSQL itself was perfectly healthy. That silently blinds exactly the alerts that matter most (WAL-archiver failures, backup staleness), because they key off primary metrics.
Removing `retentionPolicy` from the ObjectStore stops the recurring wedge immediately.
## Why I don't think it's barman itself
Running the equivalent delete manually with the very same endpoint and credentials completes in seconds:
```
barman-cloud-backup-delete --cloud-provider aws-s3 \
--endpoint-url http://:7070 \
--retention-policy "RECOVERY WINDOW OF 30 DAYS" \
s3://pg-backups/
```
So the hang appears to be in the plugin's runnable orchestration around it, not in the underlying barman operation. Possibly related to #770 (retention not pruning base backups), but the symptom here is a hang + metrics outage rather than a silent no-op.
## Expected
- The retention runnable should have a timeout, and a stuck/failed run should surface as a condition on the ObjectStore instead of hanging silently.
- A stuck retention run should never be able to take the metrics endpoint down with it.
## Current workaround
`retentionPolicy` removed; retention is done by a monthly manual CronJob running `barman-cloud-backup-delete` with the same credentials (works reliably). Happy to provide sidecar logs/goroutine dumps from a reproduction if that helps.
Guide de contribution
Ouvrir le guide de contribution
Piste de recherche
Commencez par suivre le runnable de rétention déclenché par la cadence d’ObjectStore ainsi que le handler /metrics du sidecar de l’instance, en reproduisant le blocage contre une passerelle S3 avec backend POSIX. Comparez l’orchestration du plugin avec la commande fonctionnelle barman-cloud-backup-delete. C’est terminé lorsqu’une exécution bloquée expire et fait apparaître une condition ObjectStore sans mettre /metrics hors ligne.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- go, postgresql
- Domaine
- backend, cloud, observability-sre
- Type d'issue
- Bug
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Activité
- Calme
- Clarté
- Plutôt claire
- Accessibilité débutants
- 48/100