NVIDIA / NVIDIA/nvbandwidth

Performance drop in host_to_device_memcpy_sm with large buffers

Open
#42 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
778
Forks
93
PR merge metrics
No merged PRs in 30d

Description

I am using Ubuntu 22.04 and testing the performance of an NVIDIA A100 GPU with the nvbandwidth tool. I observed that as the buffer size increases, the reported throughput decreases:

Test case: host_to_device_memcpy_sm

  • 512 MiB: 25.13 GB/s
  • 1 GiB: 25.13 GB/s
  • 10 GiB: 18.05 GB/s
  • 20 GiB: 16.50 GB/s

Below is the output:

$ ./nvbandwidth -t host_to_device_memcpy_sm -b 512
nvbandwidth Version: v0.8
Built from Git version: v0.8

CUDA Runtime Version: 12060
CUDA Driver Version: 12060
Driver Version: 560.35.05

Device 0: NVIDIA A100-SXM4-40GB (00000000:99:00)

Running host_to_device_memcpy_sm.
memcpy SM CPU(row) -> GPU(column) bandwidth (GB/s)
           0
 0     25.13

SUM host_to_device_memcpy_sm 25.13

NOTE: The reported results may not reflect the full capabilities of the platform.
Performance can vary with software drivers, hardware clocks, and system topology.
$ ./nvbandwidth -t host_to_device_memcpy_sm -b 1024
nvbandwidth Version: v0.8
Built from Git version: v0.8

CUDA Runtime Version: 12060
CUDA Driver Version: 12060
Driver Version: 560.35.05

Device 0: NVIDIA A100-SXM4-40GB (00000000:99:00)

Running host_to_device_memcpy_sm.
memcpy SM CPU(row) -> GPU(column) bandwidth (GB/s)
           0
 0     25.13

SUM host_to_device_memcpy_sm 25.13

NOTE: The reported results may not reflect the full capabilities of the platform.
Performance can vary with software drivers, hardware clocks, and system topology.
$ ./nvbandwidth -t host_to_device_memcpy_sm -b 10240
nvbandwidth Version: v0.8
Built from Git version: v0.8

CUDA Runtime Version: 12060
CUDA Driver Version: 12060
Driver Version: 560.35.05

Device 0: NVIDIA A100-SXM4-40GB (00000000:99:00)

Running host_to_device_memcpy_sm.
memcpy SM CPU(row) -> GPU(column) bandwidth (GB/s)
           0
 0     18.05

SUM host_to_device_memcpy_sm 18.05

NOTE: The reported results may not reflect the full capabilities of the platform.
Performance can vary with software drivers, hardware clocks, and system topology.
$ ./nvbandwidth -t host_to_device_memcpy_sm -b 20480
nvbandwidth Version: v0.8
Built from Git version: v0.8

CUDA Runtime Version: 12060
CUDA Driver Version: 12060
Driver Version: 560.35.05

Device 0: NVIDIA A100-SXM4-40GB (00000000:99:00)

Running host_to_device_memcpy_sm.
memcpy SM CPU(row) -> GPU(column) bandwidth (GB/s)
           0
 0     16.50

SUM host_to_device_memcpy_sm 16.50

NOTE: The reported results may not reflect the full capabilities of the platform.
Performance can vary with software drivers, hardware clocks, and system topology.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the ./nvbandwidth entry point with -t host_to_device_memcpy_sm at the reported buffer sizes on Ubuntu 22.04 and an NVIDIA A100. Compare the throughput results and investigate the host_to_device_memcpy_sm test path; done means determining the cause of the drop and documenting or correcting the behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
performance, tooling
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.