apache / apache/cloudstack

KVM NAS backup: hard NFS mount default causes VM outages when backup storage is unreachable

未关闭
#12,829 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
component:backup no-issue-activity
主要语言
Java
星标
3.1k
派生
1.4k
平均合并
6 天 19 小时
30 天内合并 PR
32

描述

## Description

When a NAS backup repository becomes temporarily unreachable, running VMs on the KVM host can be affected — including VMs that have no relationship to the backup operation. This is because the NFS mount used by `nasbackup.sh` defaults to `hard` mode, which blocks indefinitely when the NFS server is unresponsive.

## Root Cause

The `backup_repository.mount_opts` column defaults to empty. When `nasbackup.sh` calls `mount_operation()`, it mounts the backup NFS share with no special options, which defaults to NFS `hard` mode:

```bash
# nasbackup.sh mount_operation()
mount -t ${NAS_TYPE} ${NAS_ADDRESS} ${mount_point} $([[ ! -z "${MOUNT_OPTS}" ]] && echo -o ${MOUNT_OPTS})
```

With `hard` mode, any I/O operation on the NFS mount blocks indefinitely when the server is unreachable. This causes a cascade:

1. `nasbackup.sh` hangs on NFS I/O (write, sync, or umount)
2. The CloudStack agent is blocked because `nasbackup.sh` runs as a child process of the agent JVM
3. The blocked agent cannot process **any** VM operations (PlugNic, Stop, Migrate, etc.) — all commands queue behind the stuck backup
4. On the host kernel level, NFS `hard` mount stalls can cause I/O waits that affect all processes, including QEMU instances for unrelated VMs
5. VMs experience I/O timeouts — Windows guests BSOD with `KERNEL_DATA_INPAGE_ERROR`, Linux guests may freeze

## Evidence (from production CloudStack 4.20)

NFS server 172.16.3.63 experienced intermittent connectivity issues. Host dmesg showed repeated cycles:
```
nfs: server 172.16.3.63 not responding, still trying
nfs: server 172.16.3.63 OK
nfs: server 172.16.3.63 not responding, still trying
nfs: server 172.16.3.63 OK
```

Impact:
- A Windows VM (`citytravelsacco`, i-2-1651-VM) on the same host **crashed with BSOD** (`KERNEL_DATA_INPAGE_ERROR`) even though its disk is on **local storage**, not NFS
- The CloudStack agent was blocked for 3+ hours by a stuck `nasbackup.sh` process, preventing all VM management operations on the host
- A NIC hot-plug operation queued for 30+ minutes waiting for the agent to become responsive

## Suggested Fix

1. **Default `mount_opts` for NAS backup repositories to `soft,timeo=50,retrans=3`** — this causes NFS operations to fail after ~15 seconds instead of blocking forever. A failed backup is far preferable to crashing production VMs.

2. **Add a timeout wrapper to `nasbackup.sh`** — if the entire backup operation exceeds a configurable duration, kill it cleanly (resume paused VM, unmount, exit with error).

3. **Document the risk** — warn administrators that empty `mount_opts` on NAS backup repositories defaults to NFS `hard` mode, which can cause host-wide I/O stalls.

Note: `soft` mount may cause backup data corruption if the NFS server recovers mid-write, but this only affects the backup copy — not the production VM. A corrupted backup can be retried; a crashed production VM cannot be un-crashed.

## Workaround

Manually set mount options on the backup repository:
```sql
UPDATE cloud.backup_repository SET mount_opts='soft,timeo=50,retrans=3' WHERE id=;
```

And update `/etc/fstab` on KVM hosts if the NFS backup share is persistently mounted:
```
172.16.3.63:/ACS /tmp/nasbackup nfs soft,timeo=50,retrans=3,_netdev 0 0
```

## Versions

- CloudStack 4.20
- NFS v4.1
- KVM hosts: Debian/Ubuntu with kernel 5.x/6.x

贡献指南

打开贡献指南

调研方向

从 nasbackup.sh 及其入口点 mount_operation() 开始,然后跟踪 backup_repository.mount_opts 的默认值以及 CloudStack agent 的进程边界。复现或审查 issue 中描述的不可访问 NFS 行为。完成要求包括:针对 hard mounts 达成一致的缓解措施、任何所需的超时或清理行为,以及对运行风险的文档说明。

由索引模型根据 Issue 内容生成。

评估

技术栈
java, linux, shell
领域
backend, infrastructure, operating-systems
Issue 类型
缺陷
难度
5/5
预计耗时
一周以上
活跃度
冷清
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。