cloudnative-pg / cloudnative-pg/plugin-barman-cloud

plugin-barman-cloud does not clean PGDATA before re-extracting the base backup on Job retry, causing disk accumulation across failed attempts

Aperta
#1,100 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Go
Stelle
191
Fork
72
Merge medio
2g 21h
PR unite (30g)
21

Descrizione

### Environment

- CloudNativePG operator: 1.26.0
- plugin-barman-cloud: v0.6.0
- PostgreSQL: 17.5
- Restore path: `bootstrap.recovery` (externalClusters + barman-cloud plugin), object storage on Azure Blob, restore PVC ~50Gi for a ~8GB source database

### What happened

A bootstrap recovery Job failed repeatedly against a backup with a WAL-archiving gap (separate report filed against cloudnative-pg/cloudnative-pg about the Job always retrying the same pinned backup). Kubernetes' default `Job.spec.backoffLimit: 6` gave it 7 total pod attempts. The first 6 attempts each:

1. Ran `barman-cloud-restore` successfully, extracting the full base backup (~8GB) into `PGDATA` on the restore PVC.
2. Started PostgreSQL in recovery, failed WAL replay, and exited.

Nothing cleaned `PGDATA` between these attempts. Each retry re-extracted the same ~8GB base backup on top of / alongside whatever the previous attempt left behind, so usage accumulated across attempts. By the 7th attempt, the ~50Gi restore PVC was full, and that attempt failed differently and much earlier:

```
ERROR: Barman cloud restore exception: [Errno 28] No space left on device
```

### Why this is a problem

This turns a clear, correctly-diagnosable failure (a WAL-archiving gap, "WAL ends before end of online backup") into a confusing, unrelated-looking failure (disk exhaustion) purely as a side effect of retrying without cleanup. Whoever investigates the final Job state sees the disk-space error, not the real cause, unless they dig through every earlier retry pod's logs individually. It also means a PVC sized correctly for the database itself is not necessarily sized correctly for `backoffLimit + 1` retries of it.

### Suggestion

Clean/wipe the target `PGDATA` directory (and pgdata-adjacent restore artifacts) at the start of each restore attempt, before `barman-cloud-restore` re-extracts the base backup — or at minimum, detect and fail fast if `PGDATA` is non-empty going into a retry, rather than silently extracting on top of leftover data from a previous failed attempt.

Happy to provide more logs/detail if useful. This was observed in a production DR setup, not a lab reproduction, so some specifics have been generalized above.

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia dal percorso di ripristino di bootstrap.recovery e dall’invocazione di barman-cloud-restore, concentrandoti su come un Kubernetes Job prepara PGDATA per un nuovo tentativo. Verifica il comportamento con ripetuti tentativi di ripristino falliti, quindi assicurati che ogni tentativo rimuova PGDATA e gli artefatti di ripristino adiacenti prima dell’estrazione, in modo che l’utilizzo del disco non si accumuli e che l’errore di ripristino originale rimanga diagnosticabile.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
azure, go, kubernetes, postgresql
Ambito
cloud, databases
Tipo di issue
Bug
Difficoltà
3/5
Tempo stimato
1-2 giorni
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
68/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.