NVIDIA-NeMo / NVIDIA-NeMo/Curator

Optimize interleaved pipeline for numpy array buffer cache

Open
#1,759 0 comments 0 reactions 1 assignee View on GitHub

@meatybobby is already working on this.

Since May 8, 2026.

enhancement
Dominant language
Python
Stars
1.8k
Forks
328
Avg merge
4d 5h
Merged PRs (30d)
30

Description

Is your feature request related to a problem? Please describe.
Mentioned in https://github.com/NVIDIA-NeMo/Curator/pull/1583#discussion_r2906878911. Currently we load image bytes to numpy array every time when we apply a filter

Describe the solution you'd like
Add a buffer cache in InterleavedBatch

Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.

Additional context
Add any other context or screenshots about the feature request here.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.