NVIDIA / NVIDIA/Personal-AI-Router

Local node's GPU/CPU/memory enrichment is never refreshed after startup

Open
#4 1 comment 0 reactions 1 assignee View on GitHub

@ckelseynv is already working on this.

Since Sep 11, 2026.

Dominant language
Go
Stars
1.4k
Forks
250
Avg merge
23h 27m
Merged PRs (30d)
1

Description

The discovery directory's enrichment for the local node is captured once and then frozen. Peers converge; self does not.

What I measured

Ubuntu node with an RTX 3060, paired with a macOS node, v0.1.1-463 from the .deb.

I loaded a 7 GB model into LM Studio with full GPU offload. nvidia-smi went from 121 MiB to 9329 MiB. node-info served the new figure within its usual cadence — curl 127.0.0.1:14318/v1/node-info returned vram_used_bytes: 9782165504 with telemetryValid: true, msSince: 280. The desktop UI's VRAM graph showed the step too, because it reads the node directly.

The discovery snapshot did not move. discovery:get-nodes kept returning vram_used_bytes: 126877696 — the value from daemon start — for as long as I watched, checked at 20 s intervals over several minutes. Restarting the broker made it correct again, until the next change.

Why

onBrowse deliberately skips our own entry (daemon.go: "Never let a browse event touch our own entry. Self is registry-driven"), so the path that re-enriches every peer on every browse never runs for self. publishSelf — the only thing that enriches self — is called at startup, on service (un)register, and after an identity/address change. refreshPeersLoop refreshes advertised addresses, the mesh, cluster identity and models on its 15 s tick, but nothing re-reads self's node-info.

Worth noting that the comment above noteNodeInfo already states the intended behaviour:

The enrichment sweep repeats every peerRefreshInterval for as long as a node is advertised

That holds for peers, not for self.

Who sees it

Only consumers of the discovery snapshot — which includes nvpair-job-scheduler, since the broker fans the same snapshot to it. The UI is unaffected, which is why this is easy to miss.

Repro

  1. Pair two nodes; note the local node's gpus[].vram_used_bytes via discovery:get-nodes.
  2. Load a large model with GPU offload on that node.
  3. curl 127.0.0.1:14318/v1/node-info — fresh.
  4. discovery:get-nodes — unchanged, indefinitely.

A fix that works

Calling d.publishSelf() on the refreshPeersLoop tick converges the directory within 20 s. Verified live in both directions, loading and unloading. It does mean one self node-updated per tick, since publishSelf emits unconditionally; the tidier alternative is an applyInfo path with a changed check, mirroring applyModels. Happy to open a PR for either shape if useful.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.