`transform_reduce_sum` regresses in CCCL 3.4 on H100
Open
- Dominant language
- C++
- Stars
- 2.5k
- Forks
- 487
- Avg merge
- 2d 7h
- Merged PRs (30d)
- 296
Description
Run: `cub_bench_transform_reduce_sum_base_T_ct__I128___OffsetT_ct__I32___Elements_io__pow2__28`
GPU: gh100_p1010_0210
CPU: x86_64
Change: -65.22%
Contributor guide
Research direction
Start by running `cub_bench_transform_reduce_sum_base_T_ct__I128___OffsetT_ct__I32___Elements_io__pow2__28` on the listed H100 and compare its result with the expected CCCL 3.4 behavior. Trace the `transform_reduce_sum` benchmark and related implementation to identify the regression, then verify that the benchmark no longer reports the 65.22% decrease.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100