[Reproducibility] Request DeepSeek-R1 H100 2P2D run artifacts and network inventory / [Reproducibility] 请求 DeepSeek-R1 H100 2P2D 运行产物和网络清单
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- github-actions, python, pytorch
- Domain
- ci-cd, machine-learning, performance
Research direction
Start by inspecting the existing DeepSeek-V4 GitHub Actions artifacts and the benchmark entry points using the sa-bench command and rendered srtctl/Slurm recipes. Done means analogous DeepSeek-R1 H100 2P2D artifacts are published, including benchmark outputs, logs, hardware/network inventory, runtime versions, and driver parameters.
Written by the indexing model from the issue text.
Description
Is your feature request related to a problem? Please describe.
I am trying to reproduce the published InferenceX DeepSeek-R1 H100 1P1D disaggregated SGLang results, especially the max-dep cases with:
- prefill: TP16 / EP1 / DP attention off / 1 worker
- decode: TP16 / EP16 / DP attention on / 1 worker
- PD disaggregation enabled
- MTP/speculative decoding enabled
- closed-loop fixed-concurrency benchmark for 1k/1k, 1k/8k, and 8k/1k
The reproduced results are directionally aligned in setup, but the decode side appears significantly faster than the published InferenceX numbers. This makes the closed-loop benchmark behave differently, especially for 8k/1k where faster decode feeds new requests back into prefill more quickly and amplifies TTFT queueing.
For example, in the 1k/1k max-dep case at concurrency 64:
| Metric | InferenceX published result | Our reproduction |
|---|---|---|
| Output throughput | 1564.8 tok/s | 2639.0 to 2939.8 tok/s depending on IB/HCA exposure |
| Mean TPOT | 36.99 ms | 19.35 to 21.60 ms |
| Mean TTFT | 1662.6 ms | 1217.0 to 1383.3 ms |
In 8k/1k max-dep at concurrency 64, the faster local decode changes the request distribution under closed-loop load. In our run, the local decode phase is much shorter, while TTFT becomes much larger. When we switch to an open-loop arrival rate around the rate implied by the closed-loop run, the TTFT becomes much closer to the published value. This suggests the discrepancy may be caused by environment/runtime differences that change decode speed, and then by closed-loop feedback amplifying the observed TTFT difference.
We are not claiming this is a correctness bug. The current evidence suggests this is likely a reproducibility/artifact gap: the public benchmark result does not include enough runtime and network information to determine whether our H100 2P2D deployment is truly aligned with the original InferenceX environment.
Describe the solution you'd like
Could you publish or attach the official DeepSeek-R1 H100 2P2D artifacts for the published max-dep runs, similar to the artifacts already available for DeepSeek-V4 runs in GitHub Actions?
The most useful artifacts would be:
-
Aggregated and raw benchmark outputs:
agg_*.json- raw
results_concurrency_*.json, if available
-
Server and worker logs:
- frontend logs
- prefill worker logs
- decode worker logs
- generated
srtctl/ Slurm commands or rendered recipes
-
Per-node hardware and network inventory:
- active IB/HCA device list, for example
ibdev2netdev,ibstat,ibv_devices, or equivalent nvidia-smi topo -mnvidia-smi -qor at least GPU clocks, power limits, and MIG state- NCCL / UCX / NIXL related environment variables, such as
NCCL_IB_HCA,NCCL_SOCKET_IFNAME,UCX_NET_DEVICES,NIXL_*, etc. - whether all HCAs were exposed to the container and which interfaces were actually used by NCCL/NIXL
- active IB/HCA device list, for example
-
Runtime version information:
- InferenceX commit SHA
- srt-slurm commit SHA
- SGLang, Dynamo, CUDA, NCCL, NIXL versions
- container image tag and digest
-
Benchmark driver details:
- exact
sa-benchcommand random_range_rationum_prompts_multnum_warmup_mult- whether the published result is closed-loop only, and whether an open-loop comparison was run
- exact
For reference, existing DeepSeek-V4 GitHub Actions runs already expose useful artifacts such as:
bmk_dsv4_...server_logs_dsv4_...gpu_metrics_dsv4_...
Example public run:
https://github.com/SemiAnalysisAI/InferenceX/actions/runs/26191083562/attempts/11
That run includes artifacts with names like:
server_logs_dsv4_1k1k_fp8_vllm_tp8-ep1-dpafalse_disagg-false_spec-mtp_conc64_h200-dgxc-slurm_11gpu_metrics_dsv4_1k1k_fp8_vllm_tp8-ep1-dpafalse_disagg-false_spec-mtp_conc64_h200-dgxc-slurm_11bmk_dsv4_1k1k_fp8_vllm_tp8-ep1-dpafalse_disagg-false_spec-mtp_conc64_h200-dgxc-slurm_11
Having the analogous DeepSeek-R1 H100 2P2D artifacts would make it much easier to determine whether the decode-speed difference is due to:
- network/HCA exposure,
- NCCL/NIXL device selection,
- GPU clocks or power limits,
- SGLang/Dynamo/runtime version differences,
- CUDA graph or speculative decoding behavior,
- or a benchmark workload/closed-loop interpretation difference.
Describe alternatives you've considered
We tried varying the number of exposed IB/HCA devices locally, including 4, 6, and 8 HCA configurations. The profiles suggest that decode effective batch size, CUDA graph usage, MTP acceptance, NCCL collective behavior, and prefill/KV-transfer pressure all matter. However, without the original run logs and network/runtime inventory, we cannot tell which differences are expected environment differences and which indicate a deployment mismatch.
Additional context
The main observation is that our reproduction can be faster on decode than the published InferenceX result. In closed-loop benchmarks, this can make the system send new long-prefill requests faster, shifting more in-flight requests toward prefill and making TTFT look worse in heavy-prefill cases. This is why we think publishing the official run artifacts would help the community reproduce and interpret the results more accurately, rather than treating this as a simple pass/fail benchmark mismatch.
中文说明
用户尝试复现 InferenceX 发布的 DeepSeek-R1 H100 2P2D 分离式 SGLang 基准测试结果(max-dep 配置),复现环境在解码侧明显更快(吞吐量高约 1.7-1.9 倍),导致闭环基准测试行为不同——更快的解码使新请求更快回流到预填充队列,放大了 TTFT 排队效应。请求发布官方运行产物,包括聚合/原始基准测试输出、服务端和工作节点日志、每节点硬件和网络清单(IB/HCA 设备、NCCL/NIXL 环境变量等)、运行时版本信息,以及基准测试驱动的具体参数,以便社区更准确地复现和解读结果。
- Dominant language
- Python
- Stars
- 1.7k
- Forks
- 303
- Avg merge
- 1d 13h
- Merged PRs (30d)
- 284
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from SemiAnalysisAI/InferenceX
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
SemiAnalysisAI/InferenceX#2125 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
SemiAnalysisAI/InferenceX#1587 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 78/100
SemiAnalysisAI/InferenceX#1369 · 3 comments ·
-
Difficulty 1/5 1-3 hours Newbie friendliness 76/100
SemiAnalysisAI/InferenceX#1359 · 1 comment ·
-
Difficulty 5/5 Over a week Newbie friendliness 30/100
SemiAnalysisAI/InferenceX#3122 · 3 comments ·
All issues in SemiAnalysisAI/InferenceX
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
bancolombia/sentinel#23 ·
-
test md OpenCI
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
integration:quickjs org:external priority:backlog topic:code-interpreter topic:middleware type:feature
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
langchain-ai/deepagents#6450 ·
-
bug client
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100