NVIDIA / NVIDIA/cccl

`transform_reduce_sum` regresses in CCCL 3.4 on H100

Open
#9,760 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
2.5k
Forks
487
Avg merge
2d 7h
Merged PRs (30d)
296

Description

Run: `cub_bench_transform_reduce_sum_base_T_ct__I128___OffsetT_ct__I32___Elements_io__pow2__28`
GPU: gh100_p1010_0210
CPU: x86_64
Change: -65.22%

Contributor guide

Open the contributing guide

Research direction

Start by running `cub_bench_transform_reduce_sum_base_T_ct__I128___OffsetT_ct__I32___Elements_io__pow2__28` on the listed H100 and compare its result with the expected CCCL 3.4 behavior. Trace the `transform_reduce_sum` benchmark and related implementation to identify the regression, then verify that the benchmark no longer reports the 65.22% decrease.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.