cloudnative-pg / cloudnative-pg/plugin-barman-cloud

plugin-barman-cloud does not clean PGDATA before re-extracting the base backup on Job retry, causing disk accumulation across failed attempts

Abierto
#1,100 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Go
Estrellas
191
Forks
72
Merge medio
1 d 16 h
PR fusionados (30 d)
18

Descripción

### Environment

- CloudNativePG operator: 1.26.0
- plugin-barman-cloud: v0.6.0
- PostgreSQL: 17.5
- Restore path: `bootstrap.recovery` (externalClusters + barman-cloud plugin), object storage on Azure Blob, restore PVC ~50Gi for a ~8GB source database

### What happened

A bootstrap recovery Job failed repeatedly against a backup with a WAL-archiving gap (separate report filed against cloudnative-pg/cloudnative-pg about the Job always retrying the same pinned backup). Kubernetes' default `Job.spec.backoffLimit: 6` gave it 7 total pod attempts. The first 6 attempts each:

1. Ran `barman-cloud-restore` successfully, extracting the full base backup (~8GB) into `PGDATA` on the restore PVC.
2. Started PostgreSQL in recovery, failed WAL replay, and exited.

Nothing cleaned `PGDATA` between these attempts. Each retry re-extracted the same ~8GB base backup on top of / alongside whatever the previous attempt left behind, so usage accumulated across attempts. By the 7th attempt, the ~50Gi restore PVC was full, and that attempt failed differently and much earlier:

```
ERROR: Barman cloud restore exception: [Errno 28] No space left on device
```

### Why this is a problem

This turns a clear, correctly-diagnosable failure (a WAL-archiving gap, "WAL ends before end of online backup") into a confusing, unrelated-looking failure (disk exhaustion) purely as a side effect of retrying without cleanup. Whoever investigates the final Job state sees the disk-space error, not the real cause, unless they dig through every earlier retry pod's logs individually. It also means a PVC sized correctly for the database itself is not necessarily sized correctly for `backoffLimit + 1` retries of it.

### Suggestion

Clean/wipe the target `PGDATA` directory (and pgdata-adjacent restore artifacts) at the start of each restore attempt, before `barman-cloud-restore` re-extracts the base backup — or at minimum, detect and fail fast if `PGDATA` is non-empty going into a retry, rather than silently extracting on top of leftover data from a previous failed attempt.

Happy to provide more logs/detail if useful. This was observed in a production DR setup, not a lab reproduction, so some specifics have been generalized above.

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Comienza por la ruta de restauración de bootstrap.recovery y la invocación de barman-cloud-restore, centrándote en cómo un Kubernetes Job prepara PGDATA para un reintento. Verifica el comportamiento con intentos de restauración fallidos repetidos y, después, asegúrate de que cada intento elimine PGDATA y los artefactos de restauración adyacentes antes de la extracción, para que el uso del disco no se acumule y el fallo de restauración original siga siendo diagnosticable.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
azure, go, kubernetes, postgresql
Área
cloud, databases
Tipo de issue
Error
Dificultad
3/5
Tiempo estimado
1-2 días
Estado de actividad
Activo
Claridad
Bastante claro
Aptitud para principiantes
68/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.