apache / apache/cloudstack

Cancelling Maintenance Mode on NFS Primary Storage fails to remount on KVM hosts

Open
#12,690 9 comments 0 reactions 0 assignees View on GitHub
component:kvm component:primary-storage Severity:Major type:bug
Dominant language
Java
Stars
3.1k
Forks
1.4k
Avg merge
6d 19h
Merged PRs (30d)
32

Description

### problem

When a KVM-based NFS Primary Storage is taken out of Maintenance Mode, the CloudStack management state transitions back to "Up," but the actual mount point is not restored on the KVM hosts. This results in a silent failure where the storage is logically available in the UI but physically inaccessible on the hypervisor.

Below is the error seen in the KVM agent log

```
2026-02-23 14:09:31,922 INFO [kvm.storage.LibvirtStorageAdaptor] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Found existing defined storage pool b96dc55f-7075-3183-bcef-a0100e328e88, using it.
2026-02-23 14:09:31,922 DEBUG [utils.script.Script] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Executing command [/bin/bash -c mountpoint -q /mnt/b96dc55f-7075-3183-bcef-a0100e328e88 ].
2026-02-23 14:09:31,926 WARN [utils.script.Script] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Execution of process [49226] for command [/bin/bash -c mountpoint -q /mnt/b96dc55f-7075-3183-bcef-a0100e328e88 ] failed.
2026-02-23 14:09:31,926 DEBUG [utils.script.Script] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Exit value of process [49226] for command [/bin/bash -c mountpoint -q /mnt/b96dc55f-7075-3183-bcef-a0100e328e88 ] is [32].
2026-02-23 14:09:31,926 WARN [utils.script.Script] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Process [49226] for command [/bin/bash -c mountpoint -q /mnt/b96dc55f-7075-3183-bcef-a0100e328e88 ] encountered the error: [32].
2026-02-23 14:09:31,926 ERROR [kvm.storage.LibvirtStorageAdaptor] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) libvirt failed to mount storage pool b96dc55f-7075-3183-bcef-a0100e328e88 at /mnt/b96dc55f-7075-3183-bcef-a0100e328e88
2026-02-23 14:09:31,928 DEBUG [cloud.agent.Agent] (AgentRequest-Handler-3:[]) (logid:5f0c1bd2) Seq 1-2670916054007415034: { Ans: , MgmtId: 32987999634307, via: 1, Ver: v1, Flags: 10, [{"com.cloud.agent.api.Answer":{"result":"false","details":"Failed to create storage pool: libvirt failed to mount storage pool b96dc55f-7075-3183-bcef-a0100e328e88 at /mnt/b96dc55f-7075-3183-bcef-a0100e328e88","wait":"0","bypassHostMaintenance":"false"}}] }
```

### versions

4.22

### The steps to reproduce the bug

1. Navigate to Infrastructure > Primary Storage.

2. Select an NFS Primary Storage and click Enable Maintenance Mode.

3. Wait for the storage state to transition to Maintenance.

4. Click Cancel Maintenance Mode.

5. Observe the status in the UI (it returns to Up) and check the mount status on the KVM host via mount -l.

### What to do about it?

Current Workaround:
The issue currently requires manual intervention on the KVM host to restore connectivity:

Restarting the libvirtd service.

Executing virsh pool-destroy to force a re-initialization.

Required Fix:
The system should automatically ensure the NFS share is properly remounted and the libvirt pool is started when Maintenance Mode is cancelled, without requiring manual host-level commands.

Contributor guide

Open the contributing guide

Research direction

Start with the KVM agent's LibvirtStorageAdaptor path shown in the log and trace what happens when NFS Primary Storage leaves Maintenance Mode. Reproduce the cancellation steps, then verify that the share is mounted and the libvirt pool is started on the KVM host without manual commands; check the mount status with mount -l.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, linux
Domain
cloud, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.