huggingface / huggingface/datasets

Dataset configuration

Open
#5,694 3 comments 2 reactions 0 assignees View on GitHub
generic discussion
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

Following discussions from https://github.com/huggingface/datasets/pull/5331

We could have something like `config.json` to define the configuration of a dataset.

```json
{
"data_dir": "data"
"data_files": {
"train": "train-[0-9][0-9][0-9][0-9]-of-[0-9][0-9][0-9][0-9][0-9]*.*"
}
}
```

we could also support a list for several configs with a 'config_name' field.

The alternative was to use YAML in the README.md.

I think it could also support a `dataset_type` field to specify which dataset builder class to use, and the other parameters would be the builder's parameters. Some parameters exist for all builders like `data_files` and `data_dir`, but some parameters are builder specific like `sep` for csv.

This format would be used in `push_to_hub` to be able to push multiple configs.

cc @huggingface/datasets

EDIT: actually we're going for the YAML approach in README.md

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.