huggingface / huggingface/fineweb-2

Fails to load dataset

Open
#9 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
264
Forks
17
PR merge metrics
No merged PRs in 30d

Description

Since no `default` dataset config is published, and I would like to iterate on diverse data, I tried:
```py
from datasets import get_dataset_config_names, load_dataset, interleave_datasets

configs = get_dataset_config_names("HuggingFaceFW/fineweb-2")
print(configs)

streams = [
load_dataset("HuggingFaceFW/fineweb-2", c, split="train", streaming=True)
for c in configs
]

# Option A: round-robin (equal mixing across languages)
ds = interleave_datasets(streams, seed=42)

# ds is now an IterableDataset; languages are naturally mixed as you iterate.
for ex in ds.take(3):
print(ex.keys())
```

This prints all configs (`['aai_Latn', 'aak_Latn', 'aau_Latn', 'aaz_Latn',....`)
and then:
> ValueError: At least one valid data file must be specified, all the data_files are invalid: {'test': [], 'train': ['hf://datasets/HuggingFaceFW/fineweb-2@af9c13333eb981300149d5ca60a8e9d659b276b9/data/abi_Latn/train/000_00000.parquet']}

Minimally:
```py
from datasets import load_dataset

ds = load_dataset("HuggingFaceFW/fineweb-2", "abi_Latn", split="train", streaming=True)
ds.take(1)
```

Works on my mac, fails on my server.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the minimal Python load_dataset example for the abi_Latn train split with streaming enabled on the server, then compare that environment with the working Mac setup. The issue is resolved when the dataset loads and ds.take(1) succeeds on the server.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.