NVIDIA / NVIDIA/cccl

[EPIC]: 1D TMA/BULK exposure via `memcpy_async`

Open
#38 2 comments 0 reactions 2 assignees Claimed by @bernhardmgruber View on GitHub
libcu++
Dominant language
C++
Stars
2.5k
Forks
487
Avg merge
2d 7h
Merged PRs (30d)
296

Description

Refactor `memcpy_async` to take advantage of 1D TMA instructions when feasible.

### Tasks
- [x] https://github.com/NVIDIA/cccl/issues/57
- [ ] https://github.com/NVIDIA/cccl/issues/58

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.