GoogleCloudPlatform / GoogleCloudPlatform/cloud-sql-proxy-operator
AuthProxyWorkload status grows without pruning for Pod selectors
- 主要语言
- Go
- 星标
- 120
- 派生
- 18
- PR 合并指标
- 30 天内没有已合并 PR
描述
## Summary
`AuthProxyWorkload` status appears to grow without bound when `spec.workloadSelector.kind: Pod` matches short-lived pods with unique names. The controller replaces an existing status entry only when `(name, namespace, kind, version)` matches; otherwise it appends a new entry. I could not find pruning of status entries for pods that no longer exist.
For high-churn workloads such as workflow/batch pods, deleted pod names therefore remain in `status.WorkloadStatus` permanently. Over time the APW object can become large enough that status updates fail with Kubernetes/etcd request-size errors, leaving reconcile status stale and preventing spec changes from being reflected in status.
## Observed impact
A Pod-selector APW matching ephemeral workflow pods accumulated thousands of stale pod entries and grew to roughly megabyte-scale status. Reconcile attempts then failed with errors of the form:
```text
etcdserver: request is too large
trying to send message larger than max (... vs. 2097152)
```
Clearing `status.WorkloadStatus` allowed reconcile to succeed again, but deleting completed pods alone did not remove the stale status entries, so the object would grow again with future pod churn.
## Expected behavior
The controller should bound APW status growth for `kind: Pod` selectors, for example by pruning status entries that no longer correspond to currently matching live workloads, compacting status for Pod selectors, or documenting/guarding against using Pod selectors for high-churn pods.
## Notes
This is most visible for `kind: Pod` selectors because pod names are often unique and short lived. The same status model is less likely to grow unbounded for stable controller objects such as Deployments or StatefulSets.
I checked releases through v1.7.10 and did not find a change that appears to address this behavior.
贡献指南
调研方向
首先跟踪为 kind: Pod 选择器构建和更新 status.WorkloadStatus 的控制器协调路径。使用短生命周期的 pod 重现该行为,并检查匹配和替换的处理方式。当过时条目受到限制或被清理,并且在 pod 高度频繁变更后状态更新仍能持续成功时,即表示完成。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- go, kubernetes
- 领域
- devops, infrastructure
- Issue 类型
- 缺陷
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 冷清
- 描述清晰度
- 基本清楚
- 新手友好度
- 45/100