VM snapshot merge on KVM can leave volume.path in DB pointing to a file that no longer exists (multi-disk VMs)
- 主要語言
- Java
- 星號
- 3.1k
- 分支
- 1.4k
- 平均合併
- 6 天 19 小時
- 30 天內合併 PR
- 32
描述
Summary:
When deleting a VM snapshot on KVM for a VM with more than one disk, each disk's data gets folded back into its real file one at a time, but CloudStack only gets a single "succeeded or failed" answer for the whole set. If an earlier disk's fold finishes for real but a later disk's then fails or times out, the whole thing is reported as failed, so CloudStack never updates its record for the disk that actually finished. Its database is left pointing to a file that no longer exists, with nothing to catch or fix this later. The VM then fails to start with "Can't find volume:", and the only current fix is to manually correct the database to match the real file.
Steps to reproduce:
1. Create a VM with two or more disks on KVM.
2. Take a VM snapshot, then delete it while the VM has enough disk activity that the merge takes a while (or induce a timeout/communication failure partway through the multi-disk merge).
3. If one disk's merge completes on the host before another disk's merge fails/times out, the completed disk's volumes.path is left stale.
4. Attempt to start the VM. It fails looking for the old file.
Environment where this was observed: KVM, disk-only VM snapshots, VM with 2 disks (ROOT + DATA), primary storage on NFS.
貢獻指南
研究方向
先追蹤 KVM 在多磁碟合併中的快照刪除路徑,以及其彙總結果如何更新磁碟區路徑;該 issue 未提供檔案名稱或測試名稱。重現一次部分合併失敗,然後驗證已完成的磁碟仍保留有效的資料庫路徑,且 VM 無需手動修復資料庫即可啟動。
由索引模型根據 Issue 內容生成。
評估
- 領域
- cloud, infrastructure
- Issue 類型
- 缺陷
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 活躍度
- 活躍
- 描述清晰度
- 基本清楚
- 新手友好度
- 48/100