[FEA]: Ensure that PSTL algorithms use the environment based overloads of CUB device APIs
- Dominant language
- C++
- Stars
- 2.5k
- Forks
- 486
- Avg merge
- 2d 6h
- Merged PRs (30d)
- 295
Description
### Is this a duplicate?
- [x] I confirmed there appear to be no [duplicate issues](https://github.com/NVIDIA/cccl/issues) for this request and that I agree to the [Code of Conduct](CODE_OF_CONDUCT.md)
### Area
Thrust
### Is your feature request related to a problem? Please describe.
We want to leverage the tuning infrastructure that CUB offers for the PSTL algorithms.
For that we cannot use the stream based overloads, but need to pass the execution policy as an environment
### Describe the solution you'd like
I would love to directly call into an enviroment based overload that also takes a memory so that we can avoid multiple allocations in case we need to provide device storage for the algorithm result.
### Describe alternatives you've considered
_No response_
### Additional context
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.