NVIDIA / NVIDIA/NemoClaw

[DGX Station] Add protected Express qualification and activation evidence

Open
#8,382 1 comment 0 reactions 1 assignee Claimed by @prekshivyas View on GitHub
area: e2e area: inference area: install area: onboarding enhancement platform: dgx-station
Dominant language
TypeScript
Stars
22.5k
Forks
3.1k
Avg merge
1d 1h
Merged PRs (30d)
715

Description

## Summary

Create protected, reproducible DGX Station Express qualification and activation evidence for every advertised host profile, topology, serving profile, and supported agent.

Parent Epic: #8379

DGX Station is currently Tested with limitations for qualified single-Station GB300 profiles. Dedicated CI is absent, no-OTA DGX OS `7.6.x` still needs full Express E2E, and dual Station remains Deferred.

## Required matrix

### Host profiles

- Generic Ubuntu 24.04 ARM64.
- Accepted OTA-form DGX OS releases.
- Accepted no-OTA DGX OS `7.6.x`.
- Qualified Colossus BaseOS and AI Developer Tools images.

### Topologies and profiles

- Single-Station Nemotron Ultra default.
- Single-Station DeepSeek explicit profile.
- Automatic dual-Station Ultra qualification.
- Safe implicit no-peer fallback and explicit-peer fail-before-effects.

### Lifecycle scenarios

- Clean install and cached install.
- Same-revision rerun.
- Reboot and Docker-group login resume.
- Existing vLLM choice and continuation.
- Non-empty Docker host action-safety boundaries from #7153.
- Failure rollback, recovery, rebuild, uninstall, and exact cleanup.
- Pre-existing OpenShell version/coexistence and multi-gateway ownership.

### Functional scenarios

- Direct API chat and automatic structured tool call.
- OpenClaw chat and unhinted create/read-file tool task.
- Every additional agent advertised for Station Express.
- Profile/provider switch with route and in-sandbox binding reconciliation.

## Evidence contract

Each run must bind the NemoClaw revision, host image, driver, OpenShell version, model revision, image digest, recipe/preset/topology digests, full serving command, startup/download time, throughput, latency, memory/KV/CUDA-graph data, kernel/NVIDIA events, cleanup receipt, and baseline restoration.

## Acceptance criteria

- [ ] A protected lane or approved physical-run process owns every matrix row.
- [ ] Exact pass/fail evidence is retained for each advertised host family.
- [ ] #8093, #7898, #7899, #7900, #7994, #7105, #7011, and #7083 have explicit qualification outcomes.
- [ ] Relevant `NVRM`, `Xid`, `GSP`, `UVM`, `NV_ERR`, ECC, parser, route, and tool-call failures fail the run.
- [ ] Unrelated stopped containers and pre-existing host state are proven unchanged where preservation is required.
- [ ] Dual-Station security and trust boundaries are validated on the exact two-rail topology.
- [ ] Failed runs leave bounded resumable state or restore the previous healthy baseline.
- [ ] Platform-matrix promotion is generated only from accepted evidence.
- [ ] Single Station stays Tested with limitations and dual Station stays Deferred until their separate activation gates pass.

## Related

- Epic #8379
- #7153
- #8093
- #7898
- #7899
- #7900
- #7994
- #7105
- #7011
- #7083

Signed-off-by: Prekshi Vyas

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.