apache / apache/cloudstack

Feature Idea: 'Host/Cluster Waiting For Maintenance' Mode

Open
#10,019 7 comments 0 reactions 0 assignees View on GitHub
type:new-feature
Dominant language
Java
Stars
3.1k
Forks
1.4k
Avg merge
6d 19h
Merged PRs (30d)
32

Description

##### ISSUE TYPE

* Improvement Request

##### COMPONENT NAME

~~~
Host? Cluster? Im not sure
~~~

##### CLOUDSTACK VERSION

~~~
NA
~~~

##### CONFIGURATION

##### OS / ENVIRONMENT

##### SUMMARY

### **Current Capability**
CloudStack currently offers a 'Maintenance' Mode, which facilitates the live migration of all VMs from a host and removes the host from the cluster for maintenance.

### **Proposed Feature: "Waiting for Maintenance" Mode**
The proposed "Waiting for Maintenance" Mode introduces a preparatory state that addresses scenarios where live migration is impractical or impossible. This feature would enable gradual decommissioning or maintenance while avoiding service disruption.

### **General Idea of How It Might Work:**

**1. **Operator Responsibilities:****
- Customer communication and notification will be managed entirely by the cloud company, outside of CloudStack. This is to inform customers that they are given a time window to voluntarily restart their VMs before the cut off date.

**2. CloudStack Responsibilities:**
- Block the creation of new VMs to the host/cluster marked as 'Waiting For Maintenance'
- Ensure restarted VMs are relocated to clusters with matching host tags.

_This is actually a similar process as how AWS Cloud does it: https://aws.amazon.com/maintenance-help/_

### **Use Cases**
**Scenario 1: Decommissioning an Old Compute Cluster**

Problem:
- Legacy clusters with outdated CPU architectures cannot perform live migration due to compatibility issues
(e.g., VM freezing during migration causing downtime).
- Existing VMs must restart to migrate to a new cluster with compatible architectures.
- The old cluster remains active, risking the placement of new VMs and hindering decommissioning.

**Scenario 2: Maintenance of GPU Clusters with GPU Passthrough**

Problem:
- GPU passthrough prevents live migration, unlike vGPU setups that allow seamless migration.
- Downtime-free maintenance is not feasible, requiring customer cooperation to restart affected VMs.

##### STEPS TO REPRODUCE

~~~
NA
~~~

##### EXPECTED RESULTS

~~~
Refer to Above
~~~

##### ACTUAL RESULTS

~~~
Not able to facilitate smooth decomissioning of servers for compute where live migration is not possible.
~~~

Contributor guide

Open the contributing guide

Research direction

Start by reading CloudStack’s existing host and cluster maintenance-mode behavior, then trace how VM creation, restart, migration, and host-tag placement are handled. Define the operator-visible state and its effects for both hosts and clusters, including how restarted VMs are placed; done means new VMs avoid marked resources and restarted VMs reach compatible tagged clusters without disrupting existing workloads.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
cloud, infrastructure
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.