bigscience-workshop / bigscience-workshop/biomedical
Consider enforcing canonical train/dev/test splits for bigbio schema
Open
enhancement
- Dominant language
- Python
- Stars
- 505
- Forks
- 117
- PR merge metrics
- No merged PRs in 30d
Description
Datasets with k-fold definitions (e.g., GAD) are currently cumbersome to use. Maybe consider always enforcing train/dev/test splits, similar to what BLURB did for HoC and BIOSSES. `source` schema could preserve folds for compatibilities sake.
Contributor guide
Assessment
This issue has not been assessed yet.