NVIDIA / NVIDIA/nvcf

docs: HA prerequisites — AZ labels and dedicated node pools in both AZs

Open
#990 0 comments 0 reactions 1 assignee View on GitHub

@shobham-nv is already working on this.

Since Aug 19, 2026.

Dominant language
Go
Stars
218
Forks
72
Avg merge
1d 12h
Merged PRs (30d)
427

Description

Description

Document operator prerequisites for production HA: label nodes by AZ, provision dedicated infra pools in both AZs, enable node selectors, and publish short RTO expectations.

Definition of Done

  • Docs describe topology.kubernetes.io/zone=site-a|site-b (correct Kubernetes label)
  • Docs describe pools: nvcf.nvidia.com/workload=cassandra|vault|control-plane with capacity in both AZs
  • Docs mention global.nodeSelectors.enabled: true
  • Short RTO note included (Tier-1 ~0s; NATS ≤10s; OpenBao ≤15s)
  • Clear that single-node is local/dev only, not production HA

By submitting this issue, you acknowledge that you are an assigned member of the NVCF development team and agree to follow our code of conduct and our contributing guidelines.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.