NVIDIA / NVIDIA/OpenShell

bug(examples): podman demo scripts cannot run on macOS (chmod on SPIRE socket fails over virtiofs)

Open
#3,298 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

state:triage-needed
Dominant language
Rust
Stars
8.7k
Forks
1.3k
Avg merge
2d 11h
Merged PRs (30d)
253

Description

What happens

Running the Podman SPIFFE token exchange demo on macOS with podman machine, the SPIRE server crashes during startup:

level=info  msg="Starting Server APIs" address="[::]:8081" network=tcp
level=error msg="Fatal run error" error="chmod /run/spire/server/private/api.sock: invalid argument"
level=error msg="Server crashed" error="chmod /run/spire/server/private/api.sock: invalid argument"

Every downstream component (OIDC discovery provider, agent, gateway, sandbox) then fails as a consequence, which makes the root cause hard to locate from the symptoms.

Why

podman/spire/start-server-oidc.sh bind-mounts a host directory into the container:

-v "${server_dir}:/run/spire/server:z"

SPIRE creates its API socket in that directory and then chmods it. When SPIRE_STATE_DIR lives on the macOS host, the directory reaches the VM over virtiofs, where chmod on a unix socket returns EINVAL.

podman/spire/start-agent.sh has the same pattern for the Workload API socket:

-v "${agent_dir}:/run/spire/agent:z"

podman/README.md makes no platform statement, so a macOS host reads as a supported configuration.

Workarounds tested

  • Podman named volume for /run/spire/server avoids the chmod entirely and the server stays up. It has no host path, though, and SPIRE_AGENT_SOCKET_HOST_PATH needs to be a mountable path because the gateway passes it into sandbox containers.
  • A path native to the VM (for example under /var/tmp) works as a bind mount for both server and agent, but the scripts' host-side mkdir -p and wait_for_socket then operate on the macOS filesystem rather than the one the containers use, so they create stray directories and the socket wait times out.
  • Running the scripts entirely inside the Podman machine VM works today, with SPIRE_STATE_DIR on a VM-native path. This is what we ended up doing.

Suggested fix

Either of:

  • Document the demo as requiring a Linux host, which is the cheaper option and sets expectations correctly.
  • Place SPIRE state on a filesystem native to the container runtime and wait for readiness via podman exec inside the container rather than polling a host path. That would make the demo work unmodified on macOS.

Environment

  • macOS 15 (Darwin 25.6.0), Podman 6.1.1
  • Podman machine: Fedora CoreOS 44, kernel 7.0.11 aarch64
  • SPIRE images: ghcr.io/spiffe/spire-server:1.12.4, ghcr.io/spiffe/oidc-discovery-provider:1.12.4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with podman/spire/start-server-oidc.sh, podman/spire/start-agent.sh, and podman/README.md; reproduce the failure using Podman machine on macOS and trace the host-directory mounts and socket readiness checks. Determine whether the change documents Linux-only support or makes state and readiness work across the VM boundary, then verify the server, agent, and downstream demo components start successfully.

Written by the indexing model from the issue text.

Assessment

Tech stack
macos, shell
Domain
devops, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.