[DOC]: Add how-to guide for invoking CUB algorithms from CUDA kernels
- Dominant language
- C++
- Stars
- 2.5k
- Forks
- 486
- Avg merge
- 2d 6h
- Merged PRs (30d)
- 295
Description
### Is this a duplicate?
- [x] I confirmed there appear to be no [duplicate issues](https://github.com/NVIDIA/cccl/issues) for this bug and that I agree to the [Code of Conduct](CODE_OF_CONDUCT.md)
### Is this for new documentation, or an update to existing docs?
New
### Describe the incorrect/future/missing documentation
CUB supports use cases where selected device-wide algorithms are launched from within a CUDA kernel via CUDA Dynamic Parallelism, but the documentation does not currently provide a clear how-to guide or complete example for this workflow.
We should add a CUB documentation page that explains how to invoke a CUB algorithm from device code using CDP, including the required compilation/linking setup, temporary storage handling etc.
Apart from that, we should modify Doxygen macro `cdp_class{1}` to link to this guide.
### If this is a correction, please provide a link to the incorrect documentation. If this is a new documentation request, please link to where you have looked.
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.