NVIDIA / NVIDIA/nccl

[Question]:Regarding the issue of measurement results provided by the Inspector plugin

Open
#2,139 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

question
Dominant language
C++
Stars
5.1k
Forks
1.4k
Avg merge
2h 3m
Merged PRs (30d)
2

Description

How is this issue impacting you?

No response

Share Your Debug Logs

No response

Question

When running the alltoall test through nccl-tests, specify the use of the Inspector plugin, I obtained some measurement results. But there are two questions I would like to ask:

  1. Whether the p2p_exec_time_us measured by the NCCL Inspector plugin includes the communication time between GPUs?
  2. Does p2p_exec_time_us include the latency measurement results between GPUs on different servers?
NCCL Version

v2.30.3+cuda13.0

Your platform details

No response

Error Message & Behavior

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the NCCL Inspector plugin's measurement documentation or implementation and the nccl-tests alltoall invocation mentioned in the question. Trace what p2p_exec_time_us measures for GPU communication, including GPUs on different servers. Done means both scope questions are answered clearly and the result is documented if the current materials are ambiguous.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
distributed-systems
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.