Dstack-TEE / Dstack-TEE/dstack

Harden VMM lifecycle consistency across update, reload, and removal

Offen
#767 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
bug
Vorherrschende Sprache
Rust
Sterne
544
Forks
96
Ø Merge
17 Std. 57 Min.
Gemergte PRs (30 T.)
117

Beschreibung

## Problem

VMM lifecycle operations update files, disk metadata, in-memory state, the CID pool, and supervisor processes across multiple steps. These steps are not consistently validated, serialized, or rolled back.

Review of #766 exposed the following independent issues. They are not part of the no-TEE feature and should be addressed separately.

## Findings

- `UpdateVm` writes the compose file, encrypted environment, and user config before later resource and manifest operations complete. A later failure leaves a partial update.
- Disk resize occurs before manifest persistence and VM reload. A later failure can leave the qcow2 virtual size inconsistent with the manifest.
- Manifest writes use a direct file write, so interruption can leave a truncated manifest.
- `storage_fs` is derived from the image command line and app compose on each load. Changing either can change the expected filesystem for an existing disk.
- Start, stop, update, reload, and removal are not serialized per VM. Concurrent operations can observe or overwrite intermediate state.
- Reload rebuilds CID occupancy from running supervisor processes but does not preserve CIDs owned by stopped in-memory VMs. A new VM can reuse an existing CID. Supervisor processes excluded by annotation parsing may also fail to reserve their CIDs.
- User removal awaits port-forward cleanup after writing `.removing` but before spawning background cleanup. Cancellation of the RPC future can leave removal marked but not progressing until reload or restart.
- Orphan cleanup and reload can race: cleanup may finalize after reload has recreated state for the same VM and CID.

The existing `.removing` marker already provides restart recovery and should remain the durable source of removal intent.

## Suggested direction

- Validate the complete update before the first mutation.
- Define transactional file and disk update behavior, including rollback or an ordering that cannot expose inconsistent state.
- Serialize lifecycle operations per VM and define lock ordering for reload and CID allocation.
- Rebuild CID ownership from both supervisor state and loaded stopped VMs, rejecting ownership conflicts.
- Spawn removal cleanup before the RPC can be cancelled, while preserving `.removing` recovery.
- Add failure-injection and concurrent-operation tests for partial writes, disk resize failure, CID reuse, reload/removal races, and process restart.

## Context

These findings came from review of #766. The experimental fixes were removed from that PR to keep it scoped to development-only no-TEE support.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Start by locating the UpdateVm, reload, removal, lifecycle serialization, manifest, disk-resize, and CID-allocation entry points described in the issue. Trace their mutation and recovery paths first, then define the consistency behavior and add failure-injection and concurrent-operation tests covering partial writes, CID reuse, reload/removal races, and restart recovery.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
rust
Bereich
backend, infrastructure
Issue-Typ
Refactoring
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Ruhig
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
30/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.