Use of cuda::std::atomic is surprising
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 2.4k
- Forks
- 270
- Avg merge
- 3d 6h
- Merged PRs (30d)
- 39
Description
The code switches its whole atomics implementation to the one in CUDA if the <cuda/std/atomic> header is available. This can already be the case if CUDA is just installed globally or if the application uses CUDA in different parts and not for stdexec.
The problem is that on Windows, the cuda implementation is worse, for instance cuda::std::atomic<T>::wait seems to fall back to polling, instead of WaitOnAddress which leads to very high latencies (on the order of the scheduler tick of ~15 ms).
This is quite surprising and it would be good if it did not happen at all. Failing that it would be good if the switch were more explicit (maybe opt-in and fail if cuda atomics are needed for something?) and failing that it would be good if there was a way to easily disable the switch to cuda atomics.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by locating the atomics-selection logic that switches to <cuda/std/atomic> when the header is available, then inspect cuda::std::atomic::wait behavior on Windows versus WaitOnAddress. Define the chosen policy for avoiding or explicitly controlling CUDA atomics, and verify that globally installed CUDA no longer causes surprising selection or high-latency waits.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100