NVIDIA-NeMo / NVIDIA-NeMo/Curator
Optimize interleaved pipeline for numpy array buffer cache
Open
@meatybobby is already working on this.
Since May 8, 2026.
enhancement
- Dominant language
- Python
- Stars
- 1.8k
- Forks
- 328
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 30
Description
Is your feature request related to a problem? Please describe.
Mentioned in https://github.com/NVIDIA-NeMo/Curator/pull/1583#discussion_r2906878911. Currently we load image bytes to numpy array every time when we apply a filter
Describe the solution you'd like
Add a buffer cache in InterleavedBatch
Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.
Additional context
Add any other context or screenshots about the feature request here.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.