NVIDIA / NVIDIA/cccl

[BUG]: `cuda::launch` doesn't work correctly with kernels compiled with `.blocksareclusters`

Open
#8,327 0 comments 0 reactions 2 assignees Claimed by @pciolkosz View on GitHub
Dominant language
C++
Stars
2.5k
Forks
486
Avg merge
2d 6h
Merged PRs (30d)
295

Description

### Is this a duplicate?

- [x] I confirmed there appear to be no [duplicate issues](https://github.com/NVIDIA/cccl/issues) for this bug and that I agree to the [Code of Conduct](CODE_OF_CONDUCT.md)

### Type of Bug

Silent Failure

### Component

libcu++

### Describe the bug

When `.blocksareclusters` directive is specified for a kernel (can come from `__cluster_dims__` and `__block_size__`) and the kernel is launched by `cuda::launch` with a hierarchy that contains the cluster description, too many blocks (clusters) are being launched.

### How to Reproduce

```cpp
__global__ __cluster_dims__(2) void kernel() {}

int main()
{
cuda::stream stream{cuda::device_ref{0}};

const auto config = cuda::make_config(cuda::grid_dims<2>(), cuda::cluster_dims<3>(), cuda::block_dims<4>());

// launches 6 clusters instead of 2
cuda::launch(stream, config, kernel);
stream.sync();
}
```

### Expected behavior

Should launch the right number of clusters.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.