cloudnative-pg / cloudnative-pg/plugin-barman-cloud

plugin-barman-cloud does not clean PGDATA before re-extracting the base backup on Job retry, causing disk accumulation across failed attempts

Offen
#1,100 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Go
Sterne
191
Forks
72
Ø Merge
2 T. 21 Std.
Gemergte PRs (30 T.)
21

Beschreibung

### Environment

- CloudNativePG operator: 1.26.0
- plugin-barman-cloud: v0.6.0
- PostgreSQL: 17.5
- Restore path: `bootstrap.recovery` (externalClusters + barman-cloud plugin), object storage on Azure Blob, restore PVC ~50Gi for a ~8GB source database

### What happened

A bootstrap recovery Job failed repeatedly against a backup with a WAL-archiving gap (separate report filed against cloudnative-pg/cloudnative-pg about the Job always retrying the same pinned backup). Kubernetes' default `Job.spec.backoffLimit: 6` gave it 7 total pod attempts. The first 6 attempts each:

1. Ran `barman-cloud-restore` successfully, extracting the full base backup (~8GB) into `PGDATA` on the restore PVC.
2. Started PostgreSQL in recovery, failed WAL replay, and exited.

Nothing cleaned `PGDATA` between these attempts. Each retry re-extracted the same ~8GB base backup on top of / alongside whatever the previous attempt left behind, so usage accumulated across attempts. By the 7th attempt, the ~50Gi restore PVC was full, and that attempt failed differently and much earlier:

```
ERROR: Barman cloud restore exception: [Errno 28] No space left on device
```

### Why this is a problem

This turns a clear, correctly-diagnosable failure (a WAL-archiving gap, "WAL ends before end of online backup") into a confusing, unrelated-looking failure (disk exhaustion) purely as a side effect of retrying without cleanup. Whoever investigates the final Job state sees the disk-space error, not the real cause, unless they dig through every earlier retry pod's logs individually. It also means a PVC sized correctly for the database itself is not necessarily sized correctly for `backoffLimit + 1` retries of it.

### Suggestion

Clean/wipe the target `PGDATA` directory (and pgdata-adjacent restore artifacts) at the start of each restore attempt, before `barman-cloud-restore` re-extracts the base backup — or at minimum, detect and fail fast if `PGDATA` is non-empty going into a retry, rather than silently extracting on top of leftover data from a previous failed attempt.

Happy to provide more logs/detail if useful. This was observed in a production DR setup, not a lab reproduction, so some specifics have been generalized above.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Beginnen Sie beim Wiederherstellungspfad von bootstrap.recovery und beim Aufruf von barman-cloud-restore, und konzentrieren Sie sich darauf, wie ein Kubernetes Job PGDATA für Wiederholungsversuche vorbereitet. Überprüfen Sie das Verhalten anhand wiederholter fehlgeschlagener Wiederherstellungsversuche und stellen Sie anschließend sicher, dass jeder Versuch PGDATA und angrenzende Wiederherstellungsartefakte vor der Extraktion entfernt, damit sich die Festplattennutzung nicht aufaddiert und der ursprüngliche Wiederherstellungsfehler diagnostizierbar bleibt.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
azure, go, kubernetes, postgresql
Bereich
cloud, databases
Issue-Typ
Bug
Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Aktivitätsstatus
Aktiv
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
68/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.