Dstack-TEE / Dstack-TEE/dstack

Harden VMM lifecycle consistency across update, reload, and removal

Đang mở
#767 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
bug
Ngôn ngữ chính
Rust
Star
544
Fork
96
Merge trung bình
17 giờ 57 phút
Pull request đã merge (30 ngày)
117

Mô tả

## Problem

VMM lifecycle operations update files, disk metadata, in-memory state, the CID pool, and supervisor processes across multiple steps. These steps are not consistently validated, serialized, or rolled back.

Review of #766 exposed the following independent issues. They are not part of the no-TEE feature and should be addressed separately.

## Findings

- `UpdateVm` writes the compose file, encrypted environment, and user config before later resource and manifest operations complete. A later failure leaves a partial update.
- Disk resize occurs before manifest persistence and VM reload. A later failure can leave the qcow2 virtual size inconsistent with the manifest.
- Manifest writes use a direct file write, so interruption can leave a truncated manifest.
- `storage_fs` is derived from the image command line and app compose on each load. Changing either can change the expected filesystem for an existing disk.
- Start, stop, update, reload, and removal are not serialized per VM. Concurrent operations can observe or overwrite intermediate state.
- Reload rebuilds CID occupancy from running supervisor processes but does not preserve CIDs owned by stopped in-memory VMs. A new VM can reuse an existing CID. Supervisor processes excluded by annotation parsing may also fail to reserve their CIDs.
- User removal awaits port-forward cleanup after writing `.removing` but before spawning background cleanup. Cancellation of the RPC future can leave removal marked but not progressing until reload or restart.
- Orphan cleanup and reload can race: cleanup may finalize after reload has recreated state for the same VM and CID.

The existing `.removing` marker already provides restart recovery and should remain the durable source of removal intent.

## Suggested direction

- Validate the complete update before the first mutation.
- Define transactional file and disk update behavior, including rollback or an ordering that cannot expose inconsistent state.
- Serialize lifecycle operations per VM and define lock ordering for reload and CID allocation.
- Rebuild CID ownership from both supervisor state and loaded stopped VMs, rejecting ownership conflicts.
- Spawn removal cleanup before the RPC can be cancelled, while preserving `.removing` recovery.
- Add failure-injection and concurrent-operation tests for partial writes, disk resize failure, CID reuse, reload/removal races, and process restart.

## Context

These findings came from review of #766. The experimental fixes were removed from that PR to keep it scoped to development-only no-TEE support.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Start by locating the UpdateVm, reload, removal, lifecycle serialization, manifest, disk-resize, and CID-allocation entry points described in the issue. Trace their mutation and recovery paths first, then define the consistency behavior and add failure-injection and concurrent-operation tests covering partial writes, CID reuse, reload/removal races, and restart recovery.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
rust
Lĩnh vực
backend, infrastructure
Loại issue
Tái cấu trúc
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Ít trao đổi
Độ rõ ràng
Cần làm rõ
Mức phù hợp với người mới
30/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.