mosaicml / mosaicml/streaming

Last entry in the dataset is causing "Relative sample index $x is not present" error

Open
#677 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
1.6k
Forks
206
PR merge metrics
No merged PRs in 30d

Description

Environment

  • OS: [Ubuntu 20.04]
  • Hardware (GPU, or instance type): [H100]

When I try to load a big dataset with ~thousands of shards (each shard is ~1GB), on some of those shards I get the following error:

[rank5]:   File "/home/ubuntu/xxx/.venv/lib/python3.10/site-packages/streaming/base/array.py", line 90, in __getitem__
[rank5]:     return self.get_item(at)
[rank5]:   File "/home/ubuntu/xxx/.venv/lib/python3.10/site-packages/streaming/base/dataset.py", line 1235, in get_item
[rank5]:     sample = shard[shard_sample_id]
[rank5]:   File "/home/ubuntu/xxx/.venv/lib/python3.10/site-packages/streaming/base/array.py", line 90, in __getitem__
[rank5]:     return self.get_item(at)
[rank5]:   File "/home/ubuntu/xxx/.venv/lib/python3.10/site-packages/streaming/base/format/base/reader.py", line 319, in get_item
[rank5]:     data = self.get_sample_data(idx)
[rank5]:   File "/home/ubuntu/xxx/.venv/lib/python3.10/site-packages/streaming/base/format/mds/reader.py", line 145, in get_sample_data
[rank5]:     raise IndexError(
[rank5]: IndexError: Relative sample index 5205 is not present in the shard.01000.mds file.

But when looking into the actual data files, the shards themselves seem correct (~1 Gig each, where everything can be indexed properly except the last item). Here is that shard's index:

{'column_encodings': ['str', 'jpeg', 'str', 'np16', 'uint8', 'np16'],
 'column_names': ['caption',
                  'image',
                  'key',
                  'sscd_embeddings',
                  't5_xl_embeddings',
                  'vae_256x256_latents'],
 'column_sizes': [None, None, None, None, None, None],
 'compression': None,
 'format': 'mds',
 'hashes': [],
 'raw_data': {'basename': 'shard.01000.mds', 'bytes': 1073575555, 'hashes': {}},
 'samples': 5206,
 'size_limit': 1073741824,
 'version': 2,
 'zip_data': None}

As you can see, the index says there are 5206 samples. Which makes the sample index 5205 the last item. When I read the sample index manually, I see the following values:

>>> filename = "shard.01000.mds"
>>> offset = (1 + 5205) * 4
>>> with open(filename, 'rb', 0) as fp:
...     fp.seek(offset)
...     pair = fp.read(8)
...     begin, end = np.frombuffer(pair, np.uint32)
>>> begin, end = np.frombuffer(pair, np.uint32)
1073575555
>>> end
1868767867
>>> end - begin
795192312 (invalid value)

The problem is there is nothing after 1073575555:

>>> with open(filename, 'rb', 0) as fp:
...     fp.seek(1073575555)
...     data = fp.read()
... 
1073575555
>>> data
b''

I am assuming this happened because the sample didn't fit to the size limit but still got counted towards this index (since size_limit - 1073575555 is too smoll to fit anything), somehow? In either case, this seems to be made the dataset unusable. Will try to manually fix the index but just making you aware this is a problem.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with streaming/base/format/mds/reader.py and the get_item paths in streaming/base/array.py and streaming/base/dataset.py, then inspect the shard index and final sample offsets for shard.01000.mds. Reproduce the last-entry lookup using the reported 5206-sample index and verify that the final valid sample can be indexed without raising IndexError.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.