GoogleCloudPlatform / GoogleCloudPlatform/cloud-sql-proxy-operator

AuthProxyWorkload status grows without pruning for Pod selectors

未关闭
#773 0 条评论 2 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Go
星标
120
派生
18
PR 合并指标
30 天内没有已合并 PR

描述

## Summary

`AuthProxyWorkload` status appears to grow without bound when `spec.workloadSelector.kind: Pod` matches short-lived pods with unique names. The controller replaces an existing status entry only when `(name, namespace, kind, version)` matches; otherwise it appends a new entry. I could not find pruning of status entries for pods that no longer exist.

For high-churn workloads such as workflow/batch pods, deleted pod names therefore remain in `status.WorkloadStatus` permanently. Over time the APW object can become large enough that status updates fail with Kubernetes/etcd request-size errors, leaving reconcile status stale and preventing spec changes from being reflected in status.

## Observed impact

A Pod-selector APW matching ephemeral workflow pods accumulated thousands of stale pod entries and grew to roughly megabyte-scale status. Reconcile attempts then failed with errors of the form:

```text
etcdserver: request is too large
trying to send message larger than max (... vs. 2097152)
```

Clearing `status.WorkloadStatus` allowed reconcile to succeed again, but deleting completed pods alone did not remove the stale status entries, so the object would grow again with future pod churn.

## Expected behavior

The controller should bound APW status growth for `kind: Pod` selectors, for example by pruning status entries that no longer correspond to currently matching live workloads, compacting status for Pod selectors, or documenting/guarding against using Pod selectors for high-churn pods.

## Notes

This is most visible for `kind: Pod` selectors because pod names are often unique and short lived. The same status model is less likely to grow unbounded for stable controller objects such as Deployments or StatefulSets.

I checked releases through v1.7.10 and did not find a change that appears to address this behavior.

贡献指南

打开贡献指南

调研方向

首先跟踪为 kind: Pod 选择器构建和更新 status.WorkloadStatus 的控制器协调路径。使用短生命周期的 pod 重现该行为,并检查匹配和替换的处理方式。当过时条目受到限制或被清理,并且在 pod 高度频繁变更后状态更新仍能持续成功时,即表示完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
go, kubernetes
领域
devops, infrastructure
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
冷清
描述清晰度
基本清楚
新手友好度
45/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。