bigscience-workshop / bigscience-workshop/biomedical

Consider enforcing canonical train/dev/test splits for bigbio schema

Open
#675 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
505
Forks
117
PR merge metrics
No merged PRs in 30d

Description

Datasets with k-fold definitions (e.g., GAD) are currently cumbersome to use. Maybe consider always enforcing train/dev/test splits, similar to what BLURB did for HoC and BIOSSES. `source` schema could preserve folds for compatibilities sake.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.