NVIDIA / NVIDIA/cccl

[CUB] Vectorize shuffles in `WarpReduceBatched` for <4B types and blocked output

Open
#8,509 0 comments 0 reactions 1 assignee Claimed by @pauleonix View on GitHub
Dominant language
C++
Stars
2.5k
Forks
486
Avg merge
2d 6h
Merged PRs (30d)
295

Description

The blocked output arrangement would allow us to make more efficient use of warp shuffles by exchanging multiple 1B or 2B items with a single shuffle.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.