cloudnative-pg / cloudnative-pg/plugin-barman-cloud
Add retry logic to updateRecoveryWindow to handle concurrent ObjectStore status updates
- 主要语言
- Go
- 星标
- 192
- 派生
- 75
- 平均合并
- 1 天 16 小时
- 30 天内合并 PR
- 18
描述
## Problem
When running scheduled backups with retention policies, we observe transient errors:
```
{"level":"error","msg":"Error while updating the recovery window in the ObjectStore status stanza. Skipping.","error":"Operation cannot be fulfilled on objectstores.barmancloud.cnpg.io \"cluster-name-backup\": the object has been modified; please apply your changes to the latest version and try again"}
{"level":"error","msg":"Retention policy enforcement failed","error":"Operation cannot be fulfilled on objectstores.barmancloud.cnpg.io \"cluster-name-backup\": the object has been modified; please apply your changes to the latest version and try again"}
```
## Root Cause Analysis
After investigating the plugin source code, we identified that the `updateRecoveryWindow` function in `internal/cnpgi/instance/recovery_window.go` performs a direct status update without retry logic:
```go
// recovery_window.go:40
func updateRecoveryWindow(...) error {
// ... builds status ...
return c.Status().Update(ctx, objectStore) // No retry on conflict
}
```
This function is called from two places that can run concurrently:
1. **backup.go:169** - After a backup completes successfully
2. **retention.go:66** - During periodic retention policy enforcement (default every 5 minutes)
When both operations happen close together, Kubernetes optimistic concurrency control rejects one update because the `resourceVersion` changed between read and write.
## Evidence
The same file already has a function that correctly handles this scenario:
```go
// recovery_window.go:65 - setLastFailedBackupTime
func setLastFailedBackupTime(...) error {
return retry.RetryOnConflict(retry.DefaultBackoff, func() error {
var objectStore barmancloudv1.ObjectStore
if err := c.Get(ctx, objectStoreKey, &objectStore); err != nil {
return err
}
// ... update status ...
return c.Status().Update(ctx, &objectStore)
})
}
```
The `setLastFailedBackupTime` function uses `retry.RetryOnConflict` which:
1. Gets a fresh copy of the resource before updating
2. Retries on conflict with exponential backoff
## Impact
- **Severity**: Low - backups complete successfully, status eventually updates
- **User experience**: Confusing error messages in logs
- **Frequency**: Depends on backup/retention timing overlap (we see ~2 errors per 24h)
## Proposed Fix
Apply the same retry pattern to `updateRecoveryWindow`:
```go
func updateRecoveryWindow(
ctx context.Context,
c client.Client,
backupList *catalog.Catalog,
objectStore *barmancloudv1.ObjectStore,
serverName string,
) error {
return retry.RetryOnConflict(retry.DefaultBackoff, func() error {
// Get fresh copy
var freshObjectStore barmancloudv1.ObjectStore
if err := c.Get(ctx, client.ObjectKeyFromObject(objectStore), &freshObjectStore); err != nil {
return err
}
// Build recovery window
convertTime := func(t *time.Time) *metav1.Time {
if t == nil {
return nil
}
return ptr.To(metav1.NewTime(*t))
}
recoveryWindow := freshObjectStore.Status.ServerRecoveryWindow[serverName]
recoveryWindow.FirstRecoverabilityPoint = convertTime(backupList.GetFirstRecoverabilityPoint())
recoveryWindow.LastSuccessfulBackupTime = convertTime(backupList.GetLastSuccessfulBackupTime())
if freshObjectStore.Status.ServerRecoveryWindow == nil {
freshObjectStore.Status.ServerRecoveryWindow = make(map[string]barmancloudv1.RecoveryWindow)
}
freshObjectStore.Status.ServerRecoveryWindow[serverName] = recoveryWindow
return c.Status().Update(ctx, &freshObjectStore)
})
}
```
## Environment
- Plugin version: 0.10.0
- CNPG Operator: 1.26+
- Kubernetes: 1.29+
- Object storage: AWS S3
We're happy to submit a PR if this approach looks correct.
贡献指南
调研方向
从 internal/cnpgi/instance/recovery_window.go 开始,阅读 updateRecoveryWindow 以及 setLastFailedBackupTime,然后检查其在 backup.go 和 retention.go 中的调用方。确认该更改能够处理并发的 ObjectStore 状态更新,不会出现 issue 中描述的冲突错误,同时保留恢复窗口更新。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- go, kubernetes
- 领域
- backend, infrastructure
- Issue 类型
- 缺陷
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 活跃度
- 停滞
- 描述清晰度
- 描述清楚
- 新手友好度
- 62/100