[REQ] Use 8B/16B vectorized loads in CUDA kernels for performance
Open
@c0d1f1ed is already working on this.
Since Aug 28, 2025.
feature request
- Dominant language
- Python
- Stars
- 7.1k
- Forks
- 624
- Avg merge
- 3d 17h
- Merged PRs (30d)
- 5
Description
Description
We should add a fast path to load wider types using float2/float4 loads as this can give a nice performance boost for memory-heavy kernels.
Context
We observed good speedups in Mujoco-Warp doing this, and would like to have this available for any type that is properly aligned such that users can build padded data types on top of it.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.