NVIDIA / NVIDIA/Personal-AI-Router

Local node's GPU/CPU/memory enrichment is never refreshed after startup

Đang mở
#4 1 bình luận 0 reaction 1 người được giao Xem trên GitHub

@ckelseynv đang làm issue này rồi.

Từ ngày 11/9/2026.

Ngôn ngữ chính
Go
Star
1.4k
Fork
250
Merge trung bình
23 giờ 27 phút
Pull request đã merge (30 ngày)
1

Mô tả

The discovery directory's enrichment for the local node is captured once and then frozen. Peers converge; self does not.

What I measured

Ubuntu node with an RTX 3060, paired with a macOS node, v0.1.1-463 from the .deb.

I loaded a 7 GB model into LM Studio with full GPU offload. nvidia-smi went from 121 MiB to 9329 MiB. node-info served the new figure within its usual cadence — curl 127.0.0.1:14318/v1/node-info returned vram_used_bytes: 9782165504 with telemetryValid: true, msSince: 280. The desktop UI's VRAM graph showed the step too, because it reads the node directly.

The discovery snapshot did not move. discovery:get-nodes kept returning vram_used_bytes: 126877696 — the value from daemon start — for as long as I watched, checked at 20 s intervals over several minutes. Restarting the broker made it correct again, until the next change.

Why

onBrowse deliberately skips our own entry (daemon.go: "Never let a browse event touch our own entry. Self is registry-driven"), so the path that re-enriches every peer on every browse never runs for self. publishSelf — the only thing that enriches self — is called at startup, on service (un)register, and after an identity/address change. refreshPeersLoop refreshes advertised addresses, the mesh, cluster identity and models on its 15 s tick, but nothing re-reads self's node-info.

Worth noting that the comment above noteNodeInfo already states the intended behaviour:

The enrichment sweep repeats every peerRefreshInterval for as long as a node is advertised

That holds for peers, not for self.

Who sees it

Only consumers of the discovery snapshot — which includes nvpair-job-scheduler, since the broker fans the same snapshot to it. The UI is unaffected, which is why this is easy to miss.

Repro

  1. Pair two nodes; note the local node's gpus[].vram_used_bytes via discovery:get-nodes.
  2. Load a large model with GPU offload on that node.
  3. curl 127.0.0.1:14318/v1/node-info — fresh.
  4. discovery:get-nodes — unchanged, indefinitely.

A fix that works

Calling d.publishSelf() on the refreshPeersLoop tick converges the directory within 20 s. Verified live in both directions, loading and unloading. It does mean one self node-updated per tick, since publishSelf emits unconditionally; the tidier alternative is an applyInfo path with a changed check, mirroring applyModels. Happy to open a PR for either shape if useful.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.