apache / apache/cloudstack

Cancelling Maintenance Mode on NFS Primary Storage fails to remount on KVM hosts

未關閉
#12,690 9 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
component:kvm component:primary-storage Severity:Major type:bug
主要語言
Java
星號
3.1k
分支
1.4k
平均合併
6 天 19 小時
30 天內合併 PR
32

描述

### problem

When a KVM-based NFS Primary Storage is taken out of Maintenance Mode, the CloudStack management state transitions back to "Up," but the actual mount point is not restored on the KVM hosts. This results in a silent failure where the storage is logically available in the UI but physically inaccessible on the hypervisor.

Below is the error seen in the KVM agent log

```
2026-02-23 14:09:31,922 INFO [kvm.storage.LibvirtStorageAdaptor] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Found existing defined storage pool b96dc55f-7075-3183-bcef-a0100e328e88, using it.
2026-02-23 14:09:31,922 DEBUG [utils.script.Script] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Executing command [/bin/bash -c mountpoint -q /mnt/b96dc55f-7075-3183-bcef-a0100e328e88 ].
2026-02-23 14:09:31,926 WARN [utils.script.Script] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Execution of process [49226] for command [/bin/bash -c mountpoint -q /mnt/b96dc55f-7075-3183-bcef-a0100e328e88 ] failed.
2026-02-23 14:09:31,926 DEBUG [utils.script.Script] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Exit value of process [49226] for command [/bin/bash -c mountpoint -q /mnt/b96dc55f-7075-3183-bcef-a0100e328e88 ] is [32].
2026-02-23 14:09:31,926 WARN [utils.script.Script] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Process [49226] for command [/bin/bash -c mountpoint -q /mnt/b96dc55f-7075-3183-bcef-a0100e328e88 ] encountered the error: [32].
2026-02-23 14:09:31,926 ERROR [kvm.storage.LibvirtStorageAdaptor] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) libvirt failed to mount storage pool b96dc55f-7075-3183-bcef-a0100e328e88 at /mnt/b96dc55f-7075-3183-bcef-a0100e328e88
2026-02-23 14:09:31,928 DEBUG [cloud.agent.Agent] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Seq 1-2670916054007415034: { Ans: , MgmtId: 32987999634307, via: 1, Ver: v1, Flags: 10, [{"com.cloud.agent.api.Answer":{"result":"false","details":"Failed to create storage pool: libvirt failed to mount storage pool b96dc55f-7075-3183-bcef-a0100e328e88 at /mnt/b96dc55f-7075-3183-bcef-a0100e328e88","wait":"0","bypassHostMaintenance":"false"}}] }
```

### versions

4.22

### The steps to reproduce the bug

1. Navigate to Infrastructure > Primary Storage.

2. Select an NFS Primary Storage and click Enable Maintenance Mode.

3. Wait for the storage state to transition to Maintenance.

4. Click Cancel Maintenance Mode.

5. Observe the status in the UI (it returns to Up) and check the mount status on the KVM host via mount -l.

### What to do about it?

Current Workaround:
The issue currently requires manual intervention on the KVM host to restore connectivity:

Restarting the libvirtd service.

Executing virsh pool-destroy to force a re-initialization.

Required Fix:
The system should automatically ensure the NFS share is properly remounted and the libvirt pool is started when Maintenance Mode is cancelled, without requiring manual host-level commands.

貢獻指南

開啟貢獻指南

研究方向

從日誌中顯示的 KVM agent 的 LibvirtStorageAdaptor 路徑開始,追蹤 NFS Primary Storage 離開 Maintenance Mode 時發生的情況。重現取消步驟,然後驗證共用已掛載,且 libvirt pool 在 KVM host 上無需手動指令即可啟動;使用 mount -l 檢查掛載狀態。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
java, linux
領域
cloud, infrastructure
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
停滯
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。