huggingface / huggingface/datasets

Support for identifier-based automated split construction

Open
#7,287 3 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Feature request

As far as I understand, automated construction of splits for hub datasets is currently based on either file names or directory structure ([as described here](https://huggingface.co/docs/datasets/en/repository_structure))

It would seem to be pretty useful to also allow splits to be based on identifiers of individual examples

This could be configured like
{"split_name": {"column_name": [column values in split]}}

(This in turn requires unique 'index' columns, which could be explicitly supported or just assumed to be defined appropriately by the user).

I guess a potential downside would be that shards would end up spanning different splits - is this something that can be handled somehow? Would this only affect streaming from hub?

### Motivation

The main motivation would be that all data files could be stored in a single directory, and multiple sets of splits could be generated from the same data. This is often useful for large datasets with multiple distinct sets of splits.

This could all be configured via the README.md yaml configs

### Your contribution

May be able to contribute if it seems like a good idea

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.