[Power] Power measurement roadmap

Open
#2,681 0 comments 0 reactions 1 assignee View on GitHub

@edwingao28 is already working on this.

Since Aug 25, 2026.

Assessment

This issue has not been assessed yet.

Description

Links

Roadmap

Producer
  • Single-node fixed-sequence power, NVIDIA + AMD #1558 #2323
  • Multinode disaggregated power #2437
  • Prefill/decode role watts #2553
  • Schema v2 whole-deployment energy semantics #2599
  • Single-node AgentX power #2601
  • GB200/GB300 FP4 DCGM lanes #2507
  • AMD multinode power
  • Kimi-K3 power evaluation; retain its current vLLM/external-fork clone route until separately qualified
Multinode rollout [1/5]–[5/5]
  • [1/5] Long-lived srt-slurm fork prerequisite (merged 2026-08-24) edwingao28/srt-slurm#2
  • [2/5] H200 multinode AgentX infrastructure #2683 (merged 2026-08-24)
  • [3/5] H200 GLM-5.2 + DSV4 recipe rollout #2684 (merged 2026-08-24; CI sweep 8/8 = 100%, GLM + DSV4 both publication_valid: true)
  • [4/5] GB multinode recipe rollout #2687 — merged 2026-08-25; validation run 32780220262: GB200 qwen c768 + GB200 dsv4 c16384 + GB300 qwen c5120 all required=True status=complete publication_valid=True; dsv4-GB nightly-pin lanes still blocked by rotted images (pre-existing on main)
  • [5/5] B200/B300 multinode rollout #2688 — merged 2026-08-26; validation run 32924589757: b200-dgxc + b200-nscale lanes zero failures, b200-nscale-slurm_02 vllm audit required=True complete publication_valid=True (c128/c256 windows valid), B300 both frameworks published valid power; remaining reds attributed to b300-014 enroot perms (cluster ops) and intermittent NIXL; kimik2.6 lanes blocked on runner capacity (also on main)
Multinode rollout validation
  • H200 GLM-5.2 strict validation on exact producer e5c837f: run 32324249431required: true, publication_valid: true, 32/32 GPUs, schema v2, power_valid: 1
  • H200 DSV4 optional recipe validation on producer a1b8c7af: run 32312491367
  • H200 CI full validation on e5c837f via #2684: run 32767381601 — 8/8 lanes, GLM disagg + DSV4 agg both status=complete publication_valid=True
  • Local and CI contract validation for #2683, #2684, #2687, and #2688
  • H200 DSV4 exact-pin refresh on e5c837f (optional; remains required: false)
  • H100 SA-Bench dedicated-infrastructure smoke on e5c837f if required for the fork scope
  • Native GB hardware validation for #2687
  • Native B200/B300 hardware validation for #2688
  • After each rollout merge, verify production ingestion and dashboard visibility: schema v2, power_valid scrubbing, whole-deployment J/token semantics, J/query, and role watts where applicable
Power validation series [1/4]–[4/4]
  • [1/4] Single-node energy validation #2323
  • [2/4] Multinode energy validation #2437
  • [3/4] GB200/GB300 official DCGM lane #2456
  • [4/4] AMD MI355X strict monitor lifecycle #2494
Upstream srt-slurm
  • [1/4] DCGM power artifact data layer #288
  • [2/4] Multinode DCGM collection #289 (closed; carried in the fork pin)
  • [3/4] Measurement window contract #290
  • [4/4] Offline artifact validator + telemetry docs #291 (in review; not a blocker for the long-lived fork)
Dashboard
Dominant language
Python
Stars
1.7k
Forks
303
Avg merge
1d 13h
Merged PRs (30d)
284

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from SemiAnalysisAI/InferenceX

All issues in SemiAnalysisAI/InferenceX

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.