NVIDIA / NVIDIA/OpenShell

bug: Podman driver picks the wrong host IP for the supervisor callback on multi-homed Linux hosts (regression from #2942)

Open
#3,412 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

area:compute os:linux state:validated topic:networking
Dominant language
Rust
Stars
8.7k
Forks
1.3k
Avg merge
2d 11h
Merged PRs (30d)
253

Description

User Story

As a user running the local Podman driver on a multi-homed Linux laptop (wired dock + Wi-Fi + Tailscale), I want openshell sandbox create to work without hand-editing the gateway config or bringing interfaces down.

Problem Statement

On a host with multiple network interfaces, the Podman driver and Podman disagree on which host IP host.containers.internal points to, so the in-container supervisor can't reach the gateway and provisioning fails.

  • The gateway binds its callback listener to the default-route interface plus loopback: 192.168.1.27:17670 and 127.0.0.1:17670.
  • The driver tells the supervisor to dial host.containers.internal:17670 and injects --add-host host.containers.internal:host-gateway.
  • Podman resolves host-gateway to a different interface — here the Tailscale address 100.64.0.2 (wt0, CGNAT 100.64.0.0/10), where nothing listens. Bringing Wi-Fi down didn't help; Podman then used the Tailscale IP instead of the wired one.
  • The supervisor's policy fetch fails, it exits 1, the sandbox container follows, and the sandbox enters Error.

The reported error is also misleading: ContainerExited: Container exited with code 0 refers to the sandbox container reacting to its supervisor dying. The real failure is only in the openshell-supervisor-<id> container logs.

Impact / Why This Matters

  • Consequence: sandbox creation fails outright on multi-homed hosts (dock + Wi-Fi + VPN/mesh laptops are common), and the surfaced error points at the wrong container, so it's hard to diagnose.

  • Workaround: set host_gateway_ip = "127.0.0.1" under [openshell.drivers.podman] in ~/.config/openshell/gateway.toml, then systemctl --user restart openshell-gateway. This works because the supervisor uses --network host, so loopback reaches the gateway's 127.0.0.1:17670 listener, and both host.containers.internal and 127.0.0.1 are cert SANs. The sandbox then reaches Ready.

    [openshell.drivers.podman]
    host_gateway_ip = "127.0.0.1"
    
  • Why insufficient: it relies on an undocumented field and non-obvious reasoning about host-network loopback; it isn't discoverable from the error. The default should just work on multi-homed hosts.

Confirmed Regression (#2942)

Confirmed by version bisection: this is a regression introduced by PR #2942 (RFC 0012 sandbox architecture).

  • 0.0.116 (does not contain #2942; tagged 2026-08-28): openshell sandbox create reaches a working sandbox shell with default config (no host_gateway_ip) on this multi-homed host. ✅
  • 0.0.117-dev / current main (contains #2942, merged 2026-09-16 as c1f2e7189, first released in v0.1.0-pre.2): the same default config fails with Error, and only the host_gateway_ip = "127.0.0.1" workaround makes it succeed. ❌
  • The gateway binds the same two listeners (192.168.1.27 + 127.0.0.1) on both versions, so the DefaultRouteInterface listener is not the trigger — the change is the callback path.

Mechanism:

  • #2942 split the sandbox into two containers and moved the gateway callback onto a host-networked supervisor (crates/openshell-driver-podman/src/container.rs: supervisor.netns.nsmode = "host"). Previously a single container made the callback.

  • The callback machinery predates #2942 and assumes bridge/pasta semantics: the DefaultRouteInterface listener negotiation is from #2492, and the host-gateway --add-host injection from #1637.

  • Under the same --add-host host.containers.internal:host-gateway, the network mode changes what the alias resolves to:

    Network mode host.containers.internal resolves to
    Bridge (--network openshell) — pre-#2942 style 169.254.1.2 (link-local host-gateway)
    Host netns (--network host) — #2942 supervisor non-link-local host IP (here, the Tailscale interface)
  • The supervisor even warns about this: host.openshell.internal maps to a non-link-local IP; trusted-gateway SSRF exemption disabled. The callback path was designed around a link-local host-gateway (bridge/pasta), which host networking does not provide. On a single-homed host the host-netns resolution still lands on a reachable IP; on a multi-homed host it selects the wrong interface, where nothing listens.

  • Note: the host_gateway_ip = "127.0.0.1" workaround only works because the supervisor is now host-networked (an artifact of the #2942 change) and therefore shares the host loopback.

Acceptance Criteria

  • A multi-homed Linux host reaches Ready with the Podman driver using default config (no manual host_gateway_ip).
  • The supervisor reaches the gateway regardless of which interface Podman picks for host-gateway.
  • When the supervisor can't reach the gateway, the phase error reports its failure reason, not ContainerExited: code 0.
  • Covered by a test, or the multi-homed host_gateway_ip guidance is documented.

Reproduction Steps

  1. On a Linux host whose interfaces differ from Podman's host-gateway resolution (e.g. wired dock + Wi-Fi + Tailscale), configure the Podman driver.
  2. Run openshell sandbox create --provider <any>.
  3. The sandbox enters Error during provisioning.

Confirm the mismatch:

ss -ltn | grep 17670
#   LISTEN ... 192.168.1.27:17670   (default-route iface)
#   LISTEN ...    127.0.0.1:17670

podman run --rm --network host --add-host host.containers.internal:host-gateway \
  registry.fedoraproject.org/fedora-minimal:latest getent hosts host.containers.internal
#   100.64.0.2  host.containers.internal   <-- Tailscale wt0, nothing listening there

Environment

  • OpenShell: 0.0.117-dev.167+g7e7a8d561 (development build; regression absent in 0.0.116), installed with:
    curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | OPENSHELL_VERSION=dev sh
    
  • OS / kernel: Fedora Linux 44 (Workstation) / 7.2.5-200.fc44.x86_64
  • Podman: 5.8.4 (netavark)
  • Driver: podman (rootless); gateway runs as the openshell-gateway systemd user service (native install, not containerized)
  • Network: multi-homed — wired dock 192.168.1.27, Wi-Fi 192.168.1.75 (same subnet), Tailscale wt0 in 100.64.0.0/10

Related

Related: #2540, #1952, and the macOS cases #1519 / #1634. Distinct from #1909 (containerized gateway on Fedora 44).

Logs

Supervisor container (openshell-supervisor-<id>) — the actual failure:

WARN openshell_supervisor: Policy fetch failed, retrying
Error: × Policy fetch failed after 5 attempts: failed to connect to OpenShell server

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with crates/openshell-driver-podman/src/container.rs and the supervisor callback path introduced by #2942. Reproduce the failure with the documented Podman host-network command and inspect the supervisor logs alongside the misleading phase error. Done means default configuration reaches Ready on a multi-homed Linux host, reports the supervisor's failure reason when unreachable, and has test coverage or documented guidance.

Written by the indexing model from the issue text.

Assessment

Tech stack
linux, rust
Domain
infrastructure, networking
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.