JuliaML / JuliaML/MLUtils.jl

Deterministic, parallel data iteration

Open
#68 2 comments 1 reaction 0 assignees View on GitHub
DataLoader enhancement
Dominant language
Julia
Stars
124
Forks
23
PR merge metrics
No merged PRs in 30d

Description

The parallel `eachobs` implementation is not deterministic in that observations are returned as soon as they are loaded, so they may be returned out of order. This is very performant, and fine for some use cases like training, where data should be shuffled anyway.

To give the option to have a deterministic iteration would be helpful in many use cases, though.

This could be implemented as a wrapper around an existing iterator that does the following:

- instead of iterating over `data` with the wrapped iterator, iterate over `(1:nobs(data), data)` to preserve ordering information
- collect returned observations, stripping the index
- return an observation only if all previous (by index) observations have been returned

I am unsure by how much this will affect performance and memory usage and how the interplay is with `buffersize`. Are there alternative approaches to this implementation?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.