cloudnative-pg / cloudnative-pg/plugin-barman-cloud

Retention run never completes against a POSIX-backend S3 gateway (VersityGW) and stalls the sidecar metrics endpoint

Open
#1,044 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Go
Stars
191
Forks
72
Avg merge
2d 21h
Merged PRs (30d)
21

Description

## Environment

- CloudNativePG 1.30.0, plugin-barman-cloud v0.14.0 (sidecar image `plugin-barman-cloud-sidecar:v0.14.0`)
- PostgreSQL 18.4 and 17.10 clusters (3 instances each)
- Object store: [VersityGW](https://github.com/versity/versitygw) S3 gateway with the **posix** backend, plain HTTP endpoint inside the cluster
- Base backups and WAL archiving against this endpoint work fine (nightly base from standby, continuous WAL, PITR restore drills pass)

## What happens

With any `retentionPolicy` set on the `ObjectStore` (I used `"30d"`), the catalog-maintenance/retention run starts on its ~30 minute cadence and **never completes**. There is no error surfaced anywhere — no failed condition on the ObjectStore, nothing actionable in the sidecar logs — the run just stays stuck, and a new one piles up on the next cadence.

The visible damage is on the metrics side: while a retention run is stuck, the instance sidecar's `/metrics` endpoint stops responding (scrapes time out; see companion issue about the missing deadline on the metrics path). In practice the exporter for the affected instance goes dark for hours — in my case the *primary* was unscrapeable for 6.7 hours while PostgreSQL itself was perfectly healthy. That silently blinds exactly the alerts that matter most (WAL-archiver failures, backup staleness), because they key off primary metrics.

Removing `retentionPolicy` from the ObjectStore stops the recurring wedge immediately.

## Why I don't think it's barman itself

Running the equivalent delete manually with the very same endpoint and credentials completes in seconds:

```
barman-cloud-backup-delete --cloud-provider aws-s3 \
--endpoint-url http://:7070 \
--retention-policy "RECOVERY WINDOW OF 30 DAYS" \
s3://pg-backups/
```

So the hang appears to be in the plugin's runnable orchestration around it, not in the underlying barman operation. Possibly related to #770 (retention not pruning base backups), but the symptom here is a hang + metrics outage rather than a silent no-op.

## Expected

- The retention runnable should have a timeout, and a stuck/failed run should surface as a condition on the ObjectStore instead of hanging silently.
- A stuck retention run should never be able to take the metrics endpoint down with it.

## Current workaround

`retentionPolicy` removed; retention is done by a monthly manual CronJob running `barman-cloud-backup-delete` with the same credentials (works reliably). Happy to provide sidecar logs/goroutine dumps from a reproduction if that helps.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.