cloudnative-pg / cloudnative-pg/plugin-barman-cloud

Former primary stuck at 1/2 after failover due to barman-cloud-check-wal-archive "Expected empty archive"

オープン
#828 コメント 7 件 リアクション 7 件 担当者 0 名 GitHub で見る
bug
主要言語
Go
スター
191
フォーク
72
平均マージ
2日 21時間
マージ済み PR(30日)
21

説明

### Description

After a failover/switchover on a running cluster (bootstrapped with `initdb`, **not** a recovery/restore), the `plugin-barman-cloud` sidecar on replica pods keeps failing with `"WAL archive check failed: Expected empty archive"`. This prevents the replica from becoming fully ready (1/2 containers).

The root cause is that `barman-cloud-check-wal-archive` is executed on **every WAL Archive gRPC call**, including on **replica pods**, and it fails because the S3 bucket already contains WAL files from previous timelines (which is expected after a switchover).

### Environment

- **CloudNativePG operator**: v1.28.1
- **plugin-barman-cloud**: v0.11.0
- **PostgreSQL**: 18.3 (`ghcr.io/cloudnative-pg/postgresql:18.3`)
- **Kubernetes**: v1.31
- **Object storage**: Ceph RGW (S3-compatible, via Rook)

### Cluster configuration

The cluster is bootstrapped with `initdb` (no recovery, no externalClusters):

```yaml
spec:
instances: 3
bootstrap:
initdb:
database: system
encoding: UTF8
owner: app
plugins:
- name: barman-cloud.cloudnative-pg.io
enabled: true
isWALArchiver: true
parameters:
barmanObjectName: my-object-store
```

### Steps to reproduce

1. Create a 3-instance CNPG cluster with `initdb` bootstrap and `plugin-barman-cloud` with `isWALArchiver: true`
2. Wait for all 3 pods to be 2/2 Running
3. A failover or switchover occurs (automatic or manual), changing the timeline (e.g., timeline 1 → 2 → 3)
4. After the failover, one or more replica pods get stuck at 1/2 — the `plugin-barman-cloud` sidecar blocks WAL archiving

### What happens

After the switchover, the former primary is demoted to replica. The instance manager detects leftover WAL files and explicitly triggers archiving on the demoted pod:

```
{"msg":"Detected ready WAL files in a former primary, triggering WAL archiving"}
```

This causes the `plugin-barman-cloud` sidecar to attempt WAL archiving. However, the S3 bucket legitimately contains WAL files from previous timelines (archived by this same pod when it was primary). The plugin executes `barman-cloud-check-wal-archive`, finds the bucket is not empty, and fails:

```
barman-cloud-check-wal-archive: ERROR: WAL archive check failed for server : Expected empty archive
```

plugin-barman-cloud sidecar log (replica pod):

```
{"level":"info","msg":"barman-cloud-check-wal-archive checking the first wal"}
{"level":"info","logger":"barman-cloud-check-wal-archive","msg":"ERROR: WAL archive check failed for server : Expected empty archive","pipe":"stderr"}
{"level":"error","msg":"Error invoking barman-cloud-check-wal-archive",
"options":["--endpoint-url","http://","--cloud-provider","aws-s3","s3://",""],
"exitCode":-1,"error":"exit status 1"}
```

### Expected behavior
All the cluster pods are correctly running after a switchover/failover.

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

plugin-barman-cloud プラグインと isWALArchiver を有効にした 3 インスタンスの initdb クラスターで問題を再現し、その後、以前の Primary 上で WAL Archive gRPC 呼び出しと barman-cloud-check-wal-archive の動作を追跡します。以前の timeline の WAL ファイルが存在するにもかかわらず、降格された pod がアーカイブをトリガーする Failover パスを検証します。Switchover または Failover 後に Expected empty archive エラーなしですべての pod が 2/2 に戻れば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
go, kubernetes, postgresql
領域
databases, devops
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
活発
明瞭さ
おおむね明確
初心者へのやさしさ
48/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。