cloudnative-pg / cloudnative-pg/plugin-barman-cloud
plugin-barman-cloud does not clean PGDATA before re-extracting the base backup on Job retry, causing disk accumulation across failed attempts
- 主要言語
- Go
- スター
- 192
- フォーク
- 75
- 平均マージ
- 1日 16時間
- マージ済み PR(30日)
- 18
説明
### Environment
- CloudNativePG operator: 1.26.0
- plugin-barman-cloud: v0.6.0
- PostgreSQL: 17.5
- Restore path: `bootstrap.recovery` (externalClusters + barman-cloud plugin), object storage on Azure Blob, restore PVC ~50Gi for a ~8GB source database
### What happened
A bootstrap recovery Job failed repeatedly against a backup with a WAL-archiving gap (separate report filed against cloudnative-pg/cloudnative-pg about the Job always retrying the same pinned backup). Kubernetes' default `Job.spec.backoffLimit: 6` gave it 7 total pod attempts. The first 6 attempts each:
1. Ran `barman-cloud-restore` successfully, extracting the full base backup (~8GB) into `PGDATA` on the restore PVC.
2. Started PostgreSQL in recovery, failed WAL replay, and exited.
Nothing cleaned `PGDATA` between these attempts. Each retry re-extracted the same ~8GB base backup on top of / alongside whatever the previous attempt left behind, so usage accumulated across attempts. By the 7th attempt, the ~50Gi restore PVC was full, and that attempt failed differently and much earlier:
```
ERROR: Barman cloud restore exception: [Errno 28] No space left on device
```
### Why this is a problem
This turns a clear, correctly-diagnosable failure (a WAL-archiving gap, "WAL ends before end of online backup") into a confusing, unrelated-looking failure (disk exhaustion) purely as a side effect of retrying without cleanup. Whoever investigates the final Job state sees the disk-space error, not the real cause, unless they dig through every earlier retry pod's logs individually. It also means a PVC sized correctly for the database itself is not necessarily sized correctly for `backoffLimit + 1` retries of it.
### Suggestion
Clean/wipe the target `PGDATA` directory (and pgdata-adjacent restore artifacts) at the start of each restore attempt, before `barman-cloud-restore` re-extracts the base backup — or at minimum, detect and fail fast if `PGDATA` is non-empty going into a retry, rather than silently extracting on top of leftover data from a previous failed attempt.
Happy to provide more logs/detail if useful. This was observed in a production DR setup, not a lab reproduction, so some specifics have been generalized above.
コントリビューションガイド
調査の方向性
bootstrap.recovery のリストアパスと barman-cloud-restore の呼び出しから始め、Kubernetes Job がリトライに備えて PGDATA をどのように準備するかに注目してください。リストアが繰り返し失敗する場合の動作を確認し、その後、ディスク使用量が蓄積せず、元のリストア失敗を診断可能な状態に保てるよう、各試行で抽出前に PGDATA と隣接するリストア成果物を削除することを確認してください。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- azure, go, kubernetes, postgresql
- 領域
- cloud, databases
- issue の種類
- バグ
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 活発さ
- 活発
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 68/100