e2b-dev / e2b-dev/runtime

RFC: NFSv4 support for the sandbox NFS proxy — plans and trade-offs

Open
#3,544 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Go
Stars
1.6k
Forks
438
PR merge metrics
No merged PRs in 30d

Description

Background

While investigating the NFS orphan-file issue (#3532) and auditing the port-scanner code, I traced the full NFS stack in e2b and noticed the sandbox NFS layer is pinned to NFSv3. I'd like to understand whether upgrading to NFSv4 is on the roadmap, and share the trade-offs I found in case it's useful context.


Current architecture

There are two distinct NFS hops:

Sandbox VM (envd)
   │  NFSv3 — hard-coded in nfsOptions (init.go:359)
   ▼
Orchestrator  ← go-nfs proxy (e2b fork of willscott/go-nfs)
   │  NFSv3 or v4.1 — determined by GCP Filestore tier
   ▼
GCP Filestore

The client-facing hop (orchestrator → sandbox VM) is fixed at NFSv3 via:

// packages/envd/internal/api/init.go:359
"nfsvers=3",  // nfs proxy is nfs version 3

The orchestrator uses github.com/e2b-dev/go-nfs (a fork of willscott/go-nfs) as the NFS server. willscott/go-nfs is described as "NFSv3 protocol implementation in pure Golang" — it does not implement NFSv4.


Why NFSv3 makes sense today

I found three reasons the current choice is architecturally sound:

1. pause/resume semantics

The mount options include noac,lookupcache=none with the comment:

// disable caching so that pause/resume works correctly

NFSv3 is stateless by design — each RPC is self-contained. After a VM snapshot/resume, the client retries failed operations transparently. NFSv4 maintains session state (clientid, open stateids, delegations, leases). A resumed VM would face expired leases and would need to go through the NFSv4 grace-period / RECLAIM_COMPLETE recovery protocol, which adds significant complexity to the resume path.

2. No suitable Go NFSv4 server library

There is no production-ready NFSv4 server implementation in Go. Building one from scratch (OPEN/CLOSE state machine, byte-range locking, delegation/recall, RPCSEC_GSS) is a substantial undertaking.

3. NFSv4's main wins don't apply here

NFSv4's Compound RPC and client-side caching (delegation) are its primary performance advantages. Both are negated by the noac,lookupcache=none configuration, which forces every operation to the server anyway.


Where NFSv4 could still help

Despite the above, NFSv4 has properties that could be relevant:

  • NFS silly-rename is NFSv3-specific. The orphan .nfs* file problem (#3532) exists because NFSv3's stateless model requires client-side rename-before-delete when a file has open fds. NFSv4's stateful OPEN/CLOSE model handles this at the protocol level — the server knows which files are open and can defer the final delete without renaming.
  • Single port. NFSv4 runs entirely over TCP port 2049, eliminating the separate portmapper (port 111) and mountd dependencies. The current code already runs a custom portmap server (packages/orchestrator/pkg/portmap/); NFSv4 would remove that layer.
  • Built-in security model. NFSv4 has a richer ACL and security model, which might matter for multi-tenant volume isolation (currently done at the go-nfs chroot layer).

Questions for the team

  1. Is NFSv4 support on the roadmap? Either for the orchestrator proxy layer, the Filestore-to-orchestrator mount, or both?

  2. Would replacing go-nfs with an NFSv4 implementation be considered, or is the plan to stay with NFSv3 on the proxy layer long-term?

  3. Is the noac,lookupcache=none requirement fundamental (i.e. required by pause/resume correctness), or is there a path to enabling caching that would make NFSv4's performance benefits relevant?

  4. How does the team think about the silly-rename / orphan-file problem (#3532) in light of the NFSv3 constraint? Is the cleanup-at-teardown fix (#3533) the intended long-term approach, or is removing the constraint (NFSv4) also considered?

Happy to contribute research or prototyping work if any of this is useful.

/cc @jakubno @dobrac @ValentaTomas @arkamar @tvi @tomassrnka Looking forward to your feedback.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with packages/envd/internal/api/init.go around nfsOptions and packages/orchestrator/pkg/portmap/, then review the go-nfs dependency and the pause/resume constraints described here. The issue is currently a roadmap and trade-off discussion rather than an implementation task; done would be a team decision about NFSv4 scope, caching, and the orphan-file approach.

Written by the indexing model from the issue text.

Assessment

Tech stack
gcp, go
Domain
distributed-systems, infrastructure, networking
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.