Feature Idea: 'Host/Cluster Waiting For Maintenance' Mode
- 主要言語
- Java
- スター
- 3.1k
- フォーク
- 1.4k
- 平均マージ
- 6日 19時間
- マージ済み PR(30日)
- 32
説明
##### ISSUE TYPE
* Improvement Request
##### COMPONENT NAME
~~~
Host? Cluster? Im not sure
~~~
##### CLOUDSTACK VERSION
~~~
NA
~~~
##### CONFIGURATION
##### OS / ENVIRONMENT
##### SUMMARY
### **Current Capability**
CloudStack currently offers a 'Maintenance' Mode, which facilitates the live migration of all VMs from a host and removes the host from the cluster for maintenance.
### **Proposed Feature: "Waiting for Maintenance" Mode**
The proposed "Waiting for Maintenance" Mode introduces a preparatory state that addresses scenarios where live migration is impractical or impossible. This feature would enable gradual decommissioning or maintenance while avoiding service disruption.
### **General Idea of How It Might Work:**
**1. **Operator Responsibilities:****
- Customer communication and notification will be managed entirely by the cloud company, outside of CloudStack. This is to inform customers that they are given a time window to voluntarily restart their VMs before the cut off date.
**2. CloudStack Responsibilities:**
- Block the creation of new VMs to the host/cluster marked as 'Waiting For Maintenance'
- Ensure restarted VMs are relocated to clusters with matching host tags.
_This is actually a similar process as how AWS Cloud does it: https://aws.amazon.com/maintenance-help/_
### **Use Cases**
**Scenario 1: Decommissioning an Old Compute Cluster**
Problem:
- Legacy clusters with outdated CPU architectures cannot perform live migration due to compatibility issues
(e.g., VM freezing during migration causing downtime).
- Existing VMs must restart to migrate to a new cluster with compatible architectures.
- The old cluster remains active, risking the placement of new VMs and hindering decommissioning.
**Scenario 2: Maintenance of GPU Clusters with GPU Passthrough**
Problem:
- GPU passthrough prevents live migration, unlike vGPU setups that allow seamless migration.
- Downtime-free maintenance is not feasible, requiring customer cooperation to restart affected VMs.
##### STEPS TO REPRODUCE
~~~
NA
~~~
##### EXPECTED RESULTS
~~~
Refer to Above
~~~
##### ACTUAL RESULTS
~~~
Not able to facilitate smooth decomissioning of servers for compute where live migration is not possible.
~~~
コントリビューションガイド
調査の方向性
まず、メンテナンスモードにおける CloudStack の既存のホストおよびクラスターの動作を読み、その後、VM の作成、再起動、移行、およびホストタグによる配置がどのように処理されるかを追跡します。ホストとクラスターの両方について、オペレーターから見える状態とその影響を定義し、再起動された VM の配置方法も含めます。新しい VM がマークされたリソースを避け、再起動された VM が既存のワークロードを妨げることなく、互換性のあるタグ付きクラスターに到達すれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- java
- 領域
- cloud, infrastructure
- issue の種類
- 機能追加
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 静か
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 35/100