RFC: NFSv4 support for the sandbox NFS proxy — plans and trade-offs
Nobody has claimed this yet.
- Dominant language
- Go
- Stars
- 1.6k
- Forks
- 438
- PR merge metrics
- No merged PRs in 30d
Description
Background
While investigating the NFS orphan-file issue (#3532) and auditing the port-scanner code, I traced the full NFS stack in e2b and noticed the sandbox NFS layer is pinned to NFSv3. I'd like to understand whether upgrading to NFSv4 is on the roadmap, and share the trade-offs I found in case it's useful context.
Current architecture
There are two distinct NFS hops:
Sandbox VM (envd)
│ NFSv3 — hard-coded in nfsOptions (init.go:359)
▼
Orchestrator ← go-nfs proxy (e2b fork of willscott/go-nfs)
│ NFSv3 or v4.1 — determined by GCP Filestore tier
▼
GCP Filestore
The client-facing hop (orchestrator → sandbox VM) is fixed at NFSv3 via:
// packages/envd/internal/api/init.go:359
"nfsvers=3", // nfs proxy is nfs version 3
The orchestrator uses github.com/e2b-dev/go-nfs (a fork of willscott/go-nfs) as the NFS server. willscott/go-nfs is described as "NFSv3 protocol implementation in pure Golang" — it does not implement NFSv4.
Why NFSv3 makes sense today
I found three reasons the current choice is architecturally sound:
1. pause/resume semantics
The mount options include noac,lookupcache=none with the comment:
// disable caching so that pause/resume works correctly
NFSv3 is stateless by design — each RPC is self-contained. After a VM snapshot/resume, the client retries failed operations transparently. NFSv4 maintains session state (clientid, open stateids, delegations, leases). A resumed VM would face expired leases and would need to go through the NFSv4 grace-period / RECLAIM_COMPLETE recovery protocol, which adds significant complexity to the resume path.
2. No suitable Go NFSv4 server library
There is no production-ready NFSv4 server implementation in Go. Building one from scratch (OPEN/CLOSE state machine, byte-range locking, delegation/recall, RPCSEC_GSS) is a substantial undertaking.
3. NFSv4's main wins don't apply here
NFSv4's Compound RPC and client-side caching (delegation) are its primary performance advantages. Both are negated by the noac,lookupcache=none configuration, which forces every operation to the server anyway.
Where NFSv4 could still help
Despite the above, NFSv4 has properties that could be relevant:
- NFS silly-rename is NFSv3-specific. The orphan
.nfs*file problem (#3532) exists because NFSv3's stateless model requires client-side rename-before-delete when a file has open fds. NFSv4's stateful OPEN/CLOSE model handles this at the protocol level — the server knows which files are open and can defer the final delete without renaming. - Single port. NFSv4 runs entirely over TCP port 2049, eliminating the separate portmapper (port 111) and mountd dependencies. The current code already runs a custom portmap server (
packages/orchestrator/pkg/portmap/); NFSv4 would remove that layer. - Built-in security model. NFSv4 has a richer ACL and security model, which might matter for multi-tenant volume isolation (currently done at the
go-nfschroot layer).
Questions for the team
-
Is NFSv4 support on the roadmap? Either for the orchestrator proxy layer, the Filestore-to-orchestrator mount, or both?
-
Would replacing
go-nfswith an NFSv4 implementation be considered, or is the plan to stay with NFSv3 on the proxy layer long-term? -
Is the
noac,lookupcache=nonerequirement fundamental (i.e. required by pause/resume correctness), or is there a path to enabling caching that would make NFSv4's performance benefits relevant? -
How does the team think about the silly-rename / orphan-file problem (#3532) in light of the NFSv3 constraint? Is the cleanup-at-teardown fix (#3533) the intended long-term approach, or is removing the constraint (NFSv4) also considered?
Happy to contribute research or prototyping work if any of this is useful.
/cc @jakubno @dobrac @ValentaTomas @arkamar @tvi @tomassrnka Looking forward to your feedback.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with packages/envd/internal/api/init.go around nfsOptions and packages/orchestrator/pkg/portmap/, then review the go-nfs dependency and the pause/resume constraints described here. The issue is currently a roadmap and trade-off discussion rather than an implementation task; done would be a team decision about NFSv4 scope, caching, and the orphan-file approach.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- gcp, go
- Domain
- distributed-systems, infrastructure, networking
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100