NVIDIA / NVIDIA/warp

[REQ] Use 8B/16B vectorized loads in CUDA kernels for performance

Open
#712 3 comments 0 reactions 1 assignee View on GitHub

@c0d1f1ed is already working on this.

Since Aug 28, 2025.

feature request
Dominant language
Python
Stars
7.1k
Forks
624
Avg merge
3d 17h
Merged PRs (30d)
5

Description

Description

We should add a fast path to load wider types using float2/float4 loads as this can give a nice performance boost for memory-heavy kernels.

Context

We observed good speedups in Mujoco-Warp doing this, and would like to have this available for any type that is properly aligned such that users can build padded data types on top of it.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.