[EPIC]: 1D TMA/BULK exposure via `memcpy_async`
Open
libcu++
- Dominant language
- C++
- Stars
- 2.5k
- Forks
- 487
- Avg merge
- 2d 7h
- Merged PRs (30d)
- 296
Description
Refactor `memcpy_async` to take advantage of 1D TMA instructions when feasible.
### Tasks
- [x] https://github.com/NVIDIA/cccl/issues/57
- [ ] https://github.com/NVIDIA/cccl/issues/58
Contributor guide
Assessment
This issue has not been assessed yet.