Most datasets in preparation?
Nobody has claimed this yet.
- Dominant language
- PHP
- Stars
- 755
- Forks
- 128
- PR merge metrics
- No merged PRs in 30d
Description
I just realized that of the 20000 datasets we advertise, only about 2500 are available, the rest are in preparation.
It's unclear what that means, given that they are mostly months or years old uploads.
If the semantics of "available" is "human (@joaquinvanschoren) verified" we should call it verified and also list unverified datasets by default.
This also seems to suggest that most people that uploaded something to openml never saw their dataset show up on the site, which is not great.
So we should do at least one of two things:
a) activate or decline most datasets
b) rephrase / rename what "active" means.
Right now I feel like saying that openml hosts 20000 datasets looks like it's stretching the truth (even though all the functionality might be available for "in preparation" datasets - not sure).
cc @berndbischl @janvanrijn @joaquinvanschoren @mfeurer
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the issue and its 33-comment thread to understand the unresolved meaning of dataset availability and the proposed alternatives. No files, tests, or entry points are named; the work is done only once the project has chosen and implemented a clear policy for activating, declining, or relabeling datasets.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100