cloudnative-pg / cloudnative-pg/plugin-barman-cloud

Add retry logic to updateRecoveryWindow to handle concurrent ObjectStore status updates

未关闭
#758 1 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看
bug
主要语言
Go
星标
192
派生
75
平均合并
1 天 16 小时
30 天内合并 PR
18

描述

## Problem

When running scheduled backups with retention policies, we observe transient errors:

```
{"level":"error","msg":"Error while updating the recovery window in the ObjectStore status stanza. Skipping.","error":"Operation cannot be fulfilled on objectstores.barmancloud.cnpg.io \"cluster-name-backup\": the object has been modified; please apply your changes to the latest version and try again"}

{"level":"error","msg":"Retention policy enforcement failed","error":"Operation cannot be fulfilled on objectstores.barmancloud.cnpg.io \"cluster-name-backup\": the object has been modified; please apply your changes to the latest version and try again"}
```

## Root Cause Analysis

After investigating the plugin source code, we identified that the `updateRecoveryWindow` function in `internal/cnpgi/instance/recovery_window.go` performs a direct status update without retry logic:

```go
// recovery_window.go:40
func updateRecoveryWindow(...) error {
// ... builds status ...
return c.Status().Update(ctx, objectStore) // No retry on conflict
}
```

This function is called from two places that can run concurrently:
1. **backup.go:169** - After a backup completes successfully
2. **retention.go:66** - During periodic retention policy enforcement (default every 5 minutes)

When both operations happen close together, Kubernetes optimistic concurrency control rejects one update because the `resourceVersion` changed between read and write.

## Evidence

The same file already has a function that correctly handles this scenario:

```go
// recovery_window.go:65 - setLastFailedBackupTime
func setLastFailedBackupTime(...) error {
return retry.RetryOnConflict(retry.DefaultBackoff, func() error {
var objectStore barmancloudv1.ObjectStore
if err := c.Get(ctx, objectStoreKey, &objectStore); err != nil {
return err
}
// ... update status ...
return c.Status().Update(ctx, &objectStore)
})
}
```

The `setLastFailedBackupTime` function uses `retry.RetryOnConflict` which:
1. Gets a fresh copy of the resource before updating
2. Retries on conflict with exponential backoff

## Impact

- **Severity**: Low - backups complete successfully, status eventually updates
- **User experience**: Confusing error messages in logs
- **Frequency**: Depends on backup/retention timing overlap (we see ~2 errors per 24h)

## Proposed Fix

Apply the same retry pattern to `updateRecoveryWindow`:

```go
func updateRecoveryWindow(
ctx context.Context,
c client.Client,
backupList *catalog.Catalog,
objectStore *barmancloudv1.ObjectStore,
serverName string,
) error {
return retry.RetryOnConflict(retry.DefaultBackoff, func() error {
// Get fresh copy
var freshObjectStore barmancloudv1.ObjectStore
if err := c.Get(ctx, client.ObjectKeyFromObject(objectStore), &freshObjectStore); err != nil {
return err
}

// Build recovery window
convertTime := func(t *time.Time) *metav1.Time {
if t == nil {
return nil
}
return ptr.To(metav1.NewTime(*t))
}

recoveryWindow := freshObjectStore.Status.ServerRecoveryWindow[serverName]
recoveryWindow.FirstRecoverabilityPoint = convertTime(backupList.GetFirstRecoverabilityPoint())
recoveryWindow.LastSuccessfulBackupTime = convertTime(backupList.GetLastSuccessfulBackupTime())

if freshObjectStore.Status.ServerRecoveryWindow == nil {
freshObjectStore.Status.ServerRecoveryWindow = make(map[string]barmancloudv1.RecoveryWindow)
}
freshObjectStore.Status.ServerRecoveryWindow[serverName] = recoveryWindow

return c.Status().Update(ctx, &freshObjectStore)
})
}
```

## Environment

- Plugin version: 0.10.0
- CNPG Operator: 1.26+
- Kubernetes: 1.29+
- Object storage: AWS S3

We're happy to submit a PR if this approach looks correct.

贡献指南

打开贡献指南

调研方向

从 internal/cnpgi/instance/recovery_window.go 开始,阅读 updateRecoveryWindow 以及 setLastFailedBackupTime,然后检查其在 backup.go 和 retention.go 中的调用方。确认该更改能够处理并发的 ObjectStore 状态更新,不会出现 issue 中描述的冲突错误,同时保留恢复窗口更新。

由索引模型根据 Issue 内容生成。

评估

技术栈
go, kubernetes
领域
backend, infrastructure
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
描述清楚
新手友好度
62/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。