cloudnative-pg / cloudnative-pg/plugin-barman-cloud
Retention run never completes against a POSIX-backend S3 gateway (VersityGW) and stalls the sidecar metrics endpoint
- Vorherrschende Sprache
- Go
- Sterne
- 191
- Forks
- 72
- Ø Merge
- 2 T. 21 Std.
- Gemergte PRs (30 T.)
- 21
Beschreibung
## Environment
- CloudNativePG 1.30.0, plugin-barman-cloud v0.14.0 (sidecar image `plugin-barman-cloud-sidecar:v0.14.0`)
- PostgreSQL 18.4 and 17.10 clusters (3 instances each)
- Object store: [VersityGW](https://github.com/versity/versitygw) S3 gateway with the **posix** backend, plain HTTP endpoint inside the cluster
- Base backups and WAL archiving against this endpoint work fine (nightly base from standby, continuous WAL, PITR restore drills pass)
## What happens
With any `retentionPolicy` set on the `ObjectStore` (I used `"30d"`), the catalog-maintenance/retention run starts on its ~30 minute cadence and **never completes**. There is no error surfaced anywhere — no failed condition on the ObjectStore, nothing actionable in the sidecar logs — the run just stays stuck, and a new one piles up on the next cadence.
The visible damage is on the metrics side: while a retention run is stuck, the instance sidecar's `/metrics` endpoint stops responding (scrapes time out; see companion issue about the missing deadline on the metrics path). In practice the exporter for the affected instance goes dark for hours — in my case the *primary* was unscrapeable for 6.7 hours while PostgreSQL itself was perfectly healthy. That silently blinds exactly the alerts that matter most (WAL-archiver failures, backup staleness), because they key off primary metrics.
Removing `retentionPolicy` from the ObjectStore stops the recurring wedge immediately.
## Why I don't think it's barman itself
Running the equivalent delete manually with the very same endpoint and credentials completes in seconds:
```
barman-cloud-backup-delete --cloud-provider aws-s3 \
--endpoint-url http://:7070 \
--retention-policy "RECOVERY WINDOW OF 30 DAYS" \
s3://pg-backups/
```
So the hang appears to be in the plugin's runnable orchestration around it, not in the underlying barman operation. Possibly related to #770 (retention not pruning base backups), but the symptom here is a hang + metrics outage rather than a silent no-op.
## Expected
- The retention runnable should have a timeout, and a stuck/failed run should surface as a condition on the ObjectStore instead of hanging silently.
- A stuck retention run should never be able to take the metrics endpoint down with it.
## Current workaround
`retentionPolicy` removed; retention is done by a monthly manual CronJob running `barman-cloud-backup-delete` with the same credentials (works reliably). Happy to provide sidecar logs/goroutine dumps from a reproduction if that helps.
Beitragsleitfaden
Rechercherichtung
Beginne damit, den durch die ObjectStore-Taktung ausgelösten Retention-Runnable und den /metrics-Handler des Instance-Sidecars nachzuverfolgen und den Hänger gegen ein S3-Gateway mit POSIX-Backend zu reproduzieren. Vergleiche die Orchestrierung des Plugins mit dem funktionierenden Befehl barman-cloud-backup-delete. Als erledigt gilt, wenn ein festgefahrener Lauf nach einem Timeout eine ObjectStore-Bedingung sichtbar macht, ohne /metrics offline zu nehmen.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- go, postgresql
- Bereich
- backend, cloud, observability-sre
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Ruhig
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 48/100