NVIDIA / NVIDIA/OpenShell

feat(supervisor): netlink/syscall network setup so the privileged supervisor needs no workload-image tools

Open
#3,280 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

state:triage-needed
Dominant language
Rust
Stars
8.7k
Forks
1.3k
Avg merge
2d 11h
Merged PRs (30d)
253

Description

User Story

As an OpenShell operator, I want a sandbox to start with the default generic Alpine image in proxy mode without the sandbox image having to ship iproute2/nftables, so that the #3116 Alpine default works on unmodified base images.

Problem Statement

In proxy mode the privileged supervisor sets up the sandbox network namespace by shelling out to tools resolved from the workload image: ip (netns/veth/addr/route — 32 call sites in crates/openshell-supervisor-process/src/netns/mod.rs), nsenter (9), nft (6), and dmesg (bypass monitoring). A bare Alpine image only ships busybox ip (no netns subcommand) and no nftables/iptables, so the supervisor fails at startup with Network namespace creation failed ... iproute2 is installed and the sandbox container exits.

The community base image previously provided these tools; #3116 removes that dependency by defaulting to bare Alpine, which surfaces the gap. Only setns (8 call sites) is already a direct syscall today — namespace/veth/route creation and firewall rules still spawn external binaries.

This is the networking subset of #2750 (make the privileged supervisor independent of workload-image code), scoped down so it can land as a focused change that unblocks #3116.

Impact / Why This Matters

Without this, the #3116 default (bare Alpine) cannot run in proxy mode — the enforced-egress path that Secure Agent Workspace and any default-deny deployment rely on. The current workarounds are to keep shipping iproute2/nftables inside every sandbox image (re-introducing the exact image dependency #3116 removes) or to disable proxy mode (losing egress isolation). Neither is acceptable for a generic default.

Proposed Design

Replace the privileged network setup's external-helper calls with in-process kernel interfaces, so the supervisor is self-contained:

  • Namespace + veth + addresses + routes: route netlink (rtnetlink) plus setns/unshare with FD-owned namespaces (removes the /run/netns requirement).
  • Enter namespaces: setns directly (already used for the enter path).
  • Firewall / bypass rules: nf_tables netlink instead of the nft binary.
  • Bypass monitoring: NFLOG instead of dmesg (also drops the CAP_SYSLOG requirement; see #2382).

The supervisor binary stays musl-static so it carries no dynamic loader or NSS dependency on the workload image. Behavior on base images that already ship iproute2/nftables must be unchanged.

Acceptance Criteria

  • A sandbox created from an unmodified docker.io/library/alpine:* image starts in proxy mode with no ip/nsenter/nft/dmesg executed from the workload image.
  • Network namespace, veth, addressing, and routing are created without spawning ip/nsenter.
  • Bypass-detection firewall rules are programmed without the nft binary.
  • Bypass monitoring works without dmesg / CAP_SYSLOG.
  • Verified on the Docker, rootless Podman, and Kubernetes runtimes.
  • No behavior change for base images that already ship the tools.

Alternatives Considered

  • Mount the tools from a trusted OpenShell image into the sandbox (Phase 1 of #2750): works, but drags dynamic libraries across libc boundaries (glibc tools on musl Alpine) and remains a mount/compatibility burden; the netlink approach removes the dependency entirely.
  • Require the default image to bundle iproute2/nftables: re-introduces the image dependency #3116 removes and bloats the generic default.

Scope

In: the networking subset above.

Out (remains in #2750): the Phase 3 execution boundary (deny execve after startup, unprivileged workload/SSH execution), the Podman health-check and Kubernetes PVC-seeding shells, the hostile-image test suite, and the VM guest path.

Related

  • Sub-task of #2750
  • Unblocks #3116 (proxy-mode networking on a bare Alpine default)
  • Related: #2382 (NFLOG bypass detection), #1335 (nftables migration)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading crates/openshell-supervisor-process/src/netns/mod.rs, especially the external ip, nsenter, nft, and dmesg call sites, and review related issues #2750, #2382, and #1335. Work through the proposed rtnetlink, setns/unshare, nf_tables, and NFLOG approach, then verify the acceptance criteria across Docker, rootless Podman, and Kubernetes. Done means proxy mode starts on an unmodified Alpine image without workload-image networking tools while existing-tool images retain their behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
linux, rust
Domain
networking, security
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.