openml / openml/OpenML

Most datasets in preparation?

Open
#794 33 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
PHP
Stars
755
Forks
128
PR merge metrics
No merged PRs in 30d

Description

I just realized that of the 20000 datasets we advertise, only about 2500 are available, the rest are in preparation.
It's unclear what that means, given that they are mostly months or years old uploads.

If the semantics of "available" is "human (@joaquinvanschoren) verified" we should call it verified and also list unverified datasets by default.

This also seems to suggest that most people that uploaded something to openml never saw their dataset show up on the site, which is not great.

So we should do at least one of two things:
a) activate or decline most datasets
b) rephrase / rename what "active" means.

Right now I feel like saying that openml hosts 20000 datasets looks like it's stretching the truth (even though all the functionality might be available for "in preparation" datasets - not sure).

cc @berndbischl @janvanrijn @joaquinvanschoren @mfeurer

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the issue and its 33-comment thread to understand the unresolved meaning of dataset availability and the proposed alternatives. No files, tests, or entry points are named; the work is done only once the project has chosen and implemented a clear policy for activating, declining, or relabeling datasets.

Written by the indexing model from the issue text.

Assessment

Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.