Dstack-TEE / Dstack-TEE/dstack

vmm: one-shot mode allocates CIDs outside the IdPool and is invisible to the VMM

Aberta
#997 0 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Rust
Estrelas
546
Forks
96
Merge médio
23h 40min
PRs com merge (30d)
126

Descrição

## Summary

One-shot mode (`vmm/src/one_shot.rs`) and the main VMM service allocate vsock CIDs from the same configured range through two mechanisms that do not know about each other. Nothing prevents them from picking the same CID.

## The two allocators

**Main service** — `IdPool` over `[cid_start, cid_start + cid_pool_size)`, rebuilt on reload from the supervisor's process list (`app.rs`, `reload_vms` / `reload_vms_sync`).

**One-shot** — `one_shot.rs:24-56`:

```rust
// scan `ps aux` for qemu-system-x86_64 ... guest-cid=
let mut one_shot_cid = config.cvm.cid_start;
while existing_cids.contains(&one_shot_cid) {
one_shot_cid += 1;
...
}
```

It starts at `cid_start`, avoids collisions by scraping `ps aux`, and never touches the pool.

## Why they can collide

One-shot launches QEMU directly (`cmd.status()` at the end of `run_one_shot`) rather than registering the process with the supervisor. `occupied_cids` in both reload paths is built from `supervisor.list()`, so **a one-shot VM's CID is invisible to the main service** and never gets occupied in the pool.

The blindness is one-directional:

| | sees the other's CIDs? | via |
|---|---|---|
| one-shot → main service | yes | `ps aux` finds the qemu processes |
| main service → one-shot | **no** | one-shot never reaches the supervisor |

So the main service can allocate a CID that a running one-shot VM already holds.

Two secondary issues in the same code path:

- **TOCTOU** — the `ps aux` scan and the QEMU launch are not atomic; a concurrent allocation in the window collides regardless.
- **Parsing** — CIDs are recovered by string-splitting `ps aux` output on `guest-cid=`, which is sensitive to how QEMU arguments are formatted.

## Note on #907

Before #907, `IdPool::allocate()` had an off-by-one that made it skip `cid_start` entirely, while one-shot starts *at* `cid_start`. That incidentally kept the two apart. #907 fixed the off-by-one (correctly — `one_shot.rs:48` and `crates/dstackup/src/cid.rs` both already treat the window as `[start, start+size)`), which removes the accidental separation.

**This is not a regression introduced by #907.** The protection only ever held for exactly one one-shot VM: a second one takes `cid_start + 1`, which was already inside the main pool's allocation range. The underlying problem is that the two allocators were never coordinated.

## Possible directions

1. Register one-shot processes with the supervisor so the existing pool machinery covers them.
2. Reserve a dedicated range for one-shot outside `[cid_start, cid_start + cid_pool_size)`.
3. Have one-shot allocate through `IdPool` rather than `ps aux`.

(1) seems most consistent with how the rest of the system tracks VMs, but one-shot is deliberately a lighter path, so (2) may be the cheaper fix.

## Confidence

The code paths are confirmed by reading: one-shot starts at `cid_start`, does not register with the supervisor, and both reload paths source `occupied_cids` from `supervisor.list()` only. **Not verified on hardware** — I have not observed an actual vsock CID collision, and I do not know how much one-shot mode is used in practice, which bounds how much this matters.

Found while reviewing #907.

Guia de contribuição

Abrir o guia de contribuição

Direção de pesquisa

Comece por vmm/src/one_shot.rs, especialmente por run_one_shot e pela varredura de CID; depois, acompanhe app.rs reload_vms/reload_vms_sync e crates/dstackup/src/cid.rs. Compare as abordagens de coordenação propostas e determine como os CIDs one-shot devem ser representados no mecanismo de alocação existente. Considera-se concluído quando alocações simultâneas one-shot e do serviço principal não puderem selecionar o mesmo CID.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
rust
Domínio
infrastructure
Tipo de issue
Bug
Dificuldade
5/5
Tempo estimado
Mais de uma semana
Status de atividade
Pouca atividade
Clareza
Razoavelmente clara
Facilidade para iniciantes
45/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.