NVIDIA / NVIDIA/OpenShell

Capability-free sandbox fails to start on kernels < 5.19 (RHEL 9.x / 5.14): seccomp WAIT_KILLABLE_RECV EINVAL

Open
#3,417 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

state:accepted
Dominant language
Rust
Stars
8.7k
Forks
1.3k
Avg merge
2d 11h
Merged PRs (30d)
253

Description

User Story

As an operator running OpenShell on OpenShift / RHEL 9.x nodes, I want the RFC-0012 capability-free sandbox to start on my existing fleet, so that I can adopt the new isolation model without waiting for a kernel bump.

Problem Statement

At startup the sandbox installs its seccomp notification listener with SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV, introduced in Linux 5.19. On older kernels the seccomp() call fails:

seccomp(SET_MODE_FILTER, NEW_LISTENER|WAIT_KILLABLE_RECV, …) = -1 EINVAL

The listener thread then dies (notification launcher disappeared) and the supervisor fails confirmation (sandbox confirmation evidence is incomplete or mismatched). This appears intentional — the code maps this EINVAL to "seccomp WAIT_KILLABLE_RECV is required (Linux 5.19 or newer)" — so this is as much a "should we relax it?" as a bug report.

Impact / Why This Matters

RHEL 9.x ships kernel 5.14 for its lifecycle, and OpenShift/RHCOS nodes run the RHEL kernel. So the capability-free sandbox has no working path on current OpenShift until RHEL 10 nodes. There is no in-cluster workaround (it is a kernel-feature gap, not a config/SCC issue). This blocks OpenShift adoption of RFC-0012 in the near term.

Reproduction Steps
  1. On a node with kernel < 5.19 (e.g. RHEL 9.8, 5.14.0-687.35.1.el9_8), create a capability-free sandbox (or run openshell-sandbox capability-probe in a zero-cap, no_new_privs pod).
  2. The probe fails at the seccomp notification step; strace shows the NEW_LISTENER|WAIT_KILLABLE_RECV call returning EINVAL.
Environment
  • OpenShift / RHCOS, RHEL 9.8, kernel 5.14.0-687.35.1.el9_8
  • RFC-0012 capability-free model (#2942), Kubernetes compute driver
Ruled out (measured on the same 5.14 node, zero capabilities)
  • Not CAP_SYS_ADMIN / SCC: a plain pod at zero caps installs a NEW_LISTENER fine under both RuntimeDefault and Unconfined.
  • Not the seccomp profile, user namespaces, SELinux, or struct-size mismatch (GET_NOTIF_SIZES = 80/24/64 on both kernels). The only differentiator is the WAIT_KILLABLE_RECV flag (5.19).
Proposed Design

Make WAIT_KILLABLE_RECV optional: attempt it, and on EINVAL fall back to a plain NEW_LISTENER; and stop gating launch-confirmation on the cancellation evidence (which is exactly this flag). Implications:

  • We lose the killable-receive semantics (the workload-side notify wait becomes interruptible rather than kill-only).
  • Zero-cap containment is unchanged — the listener still mediates every syscall; the supervisor keeps rejecting stale notifications via NOTIF_ID_VALID.
  • Validated: with this change a capability-free sandbox reaches Ready on OpenShift / kernel 5.14, end to end (workload + supervisor).

A branch implementing this (two small commits, off current main) is available: akram:fix/seccomp-wait-killable-fallback-mainhttps://github.com/NVIDIA/OpenShell/compare/main...akram:OpenShell:fix/seccomp-wait-killable-fallback-main . Happy to open it as a PR if the direction is acceptable.

Alternatives Considered
  • Require nodes on kernel ≥ 5.19 (RHEL 10 / newer RHCOS) — leaves current OpenShift users blocked.
  • Runtime-installed listener via linux.seccomp.listenerPath — heavier, and not needed for this specific gap.
Acceptance Criteria
  • The capability-free sandbox starts and reaches Ready on a kernel-5.14 node.
  • On kernels ≥ 5.19 behavior is unchanged (WAIT_KILLABLE_RECV still used).
  • The degraded semantics on < 5.19 are documented.
Open question for maintainers

Is the 5.19 floor a hard requirement (a load-bearing property of the isolation model), or the convenient baseline? If a fallback is acceptable, I have the branch above and can open a PR.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the capability-probe seccomp notification setup and the supervisor's launch-confirmation handling described in the issue; compare them with the proposed fallback branch. Determine whether the older-kernel behavior is acceptable, then validate that a capability-free sandbox reaches Ready on kernel 5.14 while retaining WAIT_KILLABLE_RECV on newer kernels and documenting degraded semantics.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, linux, rust
Domain
infrastructure, operating-systems, security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.