iptables_check.py health check deadlocks when iptables output exceeds the pipe buffer, leaking processes until the VR exhausts memory and self-reboots (~9h cycle)
- 主要语言
- Java
- 星标
- 3.1k
- 派生
- 1.4k
- 平均合并
- 6 天 19 小时
- 30 天内合并 PR
- 32
描述
##### ISSUE TYPE
* Bug Report
##### COMPONENT NAME
~~~
SystemVM / Virtual Router health checks (systemvm/debian/root/health_checks/iptables_check.py)
~~~
##### CLOUDSTACK VERSION
~~~
4.22.0.0 and 4.22.1.0 (management), systemvm template 4.22.0.
The affected code is unchanged on current main:
https://github.com/apache/cloudstack/blob/main/systemvm/debian/root/health_checks/iptables_check.py
~~~
##### CONFIGURATION
Redundant VPC virtual routers, KVM. Advanced zone. Default health check settings
(`router.health.checks.enabled=true`, `router.health.checks.advanced.interval=10`).
The affected VPC has ~200 port-forwarding/LB rules (iptables-save output: 77,029 bytes).
##### OS / ENVIRONMENT
KVM hosts on Rocky Linux 9.5/9.8 (qemu-kvm 9.0/10.1 — reproduced on both). Debian-based
system VM template, 512 MB RAM routers.
##### SUMMARY
`iptables_check.py` runs the fetch command with `Popen(cmd, shell=True, stdout=PIPE)` and
then calls `pout.wait()` **before** reading the pipe:
```python
fetchIpTableEntriesCmd = "iptables-save | grep " + destIp
pout = Popen(fetchIpTableEntriesCmd, shell=True, stdout=PIPE)
if pout.wait() != 0: # <-- deadlocks: wait() before the pipe is drained
...
ipTablesMatchingEntries = pout.communicate()[0].decode().strip().split('\n')
```
This is the deadlock explicitly warned about in the Python `subprocess` documentation:
when the command's output exceeds the OS pipe capacity (64 KiB on Linux), the child
blocks writing to the full pipe and the parent blocks forever in `wait()`.
On routers with large rule sets the check therefore hangs **on every run**. Because the
advanced health check fires every `router.health.checks.advanced.interval` (10 min
default), a new stuck process chain (`monitorServices.py advanced` -> `sh` ->
`iptables_check.py`) accumulates every interval — ~13 MB RSS each, roughly 85–90 MB/h.
On a 512 MB router, RAM and swap are exhausted after ~8.5–9 h of uptime. At that point
keepalived's `heartbeat` track script can no longer fork within its timeout
(`VRRP_Script(heartbeat) timed_out` -> `Entering FAULT STATE`), the router demotes,
wipes `/etc/keepalived/keepalived.conf`, and the guest reboots — taking the VPC down.
The impact is amplified on redundant VPCs: both routers of a pair are created together,
so their exhaustion clocks are synchronized and **both routers fail within minutes of
each other**, defeating the redundancy. We experienced a full-VPC outage roughly every
9 hours until the check was excluded.
##### STEPS TO REPRODUCE
~~~
1. Create a VPC with enough port-forwarding/LB rules that the health check's
"iptables-save | grep " output exceeds 64 KiB inside the router
(our affected router: iptables-save = 77,029 bytes total).
2. Leave router.health.checks.enabled=true with default advanced interval (10 min).
3. Inside the router, watch: ps -eo pid,ppid,etimes,rss,args | grep -E "monitorServices|iptables_check"
4. A new stuck chain appears every 600 s and never exits (parent and child both in
wchan do_wait; the shell child is blocked writing to the full pipe).
5. Watch "free -k": used memory climbs linearly; after ~8.5-9 h the router enters
keepalived FAULT, demotes/wipes, and reboots.
~~~
Observed accumulation (etimes exactly 600 s apart, none ever exiting):
~~~
90114 7551 /usr/bin/python /root/monitorServices.py advanced
90136 7551 /bin/sh -c ./health_checks/iptables_check.py advanced
90137 7551 /usr/bin/python ./health_checks/iptables_check.py advanced
92506 6952 /usr/bin/python /root/monitorServices.py advanced
... 6351 (repeats every 600s)
... 5751
... 5152
... 4552
~~~
Control group: three routers in other VPCs on the same platform with small rule sets
(iptables-save 1.2–9.3 KB) run the identical check to completion in <1 s and have
152-day uptimes. The only variable is iptables output size vs the 64 KiB pipe buffer.
##### EXPECTED RESULTS
~~~
The health check completes (or fails fast) regardless of iptables rule-set size.
~~~
##### ACTUAL RESULTS
~~~
The check hangs permanently on every run once the piped output exceeds 64 KiB.
Stuck process chains accumulate every advanced-check interval until the router
exhausts RAM+swap (~8.5-9 h on a 512 MB router), keepalived's track script starves,
the router demotes/wipes its keepalived config and reboots. Redundant pairs fail
near-simultaneously because their clocks are synchronized at creation, causing a
recurring full-VPC outage.
~~~
##### WORKAROUND
Set `router.health.checks.to.exclude = iptables_check.py` (global setting, dynamic).
Note: in our testing the value had to be the script filename **with** the `.py`
extension — the bare name `iptables_check` was silently ignored and the check kept
running (possibly a second, minor bug in how the exclusion list is matched).
Accumulated stuck processes must be killed manually inside affected routers.
##### SUGGESTED FIX
Do not call `wait()` while `stdout=PIPE` is unread. Use `communicate()` first and
check `returncode` afterwards, e.g.:
```python
pout = Popen(fetchIpTableEntriesCmd, shell=True, stdout=PIPE)
stdout, _ = pout.communicate()
if pout.returncode != 0:
...
ipTablesMatchingEntries = stdout.decode().strip().split('\n')
```
The same `Popen(stdout=PIPE)` + `wait()` pattern exists in other health check scripts
(e.g. `gateways_check.py`, `cpu_usage_check.py`). Their outputs are normally tiny so
they do not currently hang, but they share the same latent hazard and may be worth
fixing in the same pass.
贡献指南
调研方向
从 systemvm/debian/root/health_checks/iptables_check.py 开始,跟踪高级检查使用的 Popen 调用。使用大于管道缓冲区的 iptables-save | grep 输出运行它,然后验证检查能够完成且不会留下卡住的进程;如果扩大范围,则比较 gateways_check.py 和 cpu_usage_check.py 是否采用相同的 wait 模式。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- linux, python, shell
- 领域
- infrastructure, networking
- Issue 类型
- 缺陷
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 活跃度
- 活跃
- 描述清晰度
- 描述清楚
- 新手友好度
- 78/100