autogluon / autogluon/tabarena
[Suggestion] some potentially other useful subsets for BeyondArena
Open
- Dominant language
- Python
- Stars
- 303
- Forks
- 69
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 49
Description
These are some potentially useful subsets for the filter, as these data sets are common:
- Data sets with joint groups between train and test. There's already a subset with disjoint groups, so it would be easy to implement the opposite.
- Noisy data
- Sparse data
Contributor guide
Research direction
Start by tracing how the existing disjoint-groups subset is defined and exposed in the dataset-filtering code. Confirm the intended definitions for joint train/test groups, noisy data, and sparse data with the issue discussion; done means the requested subsets are available through the filter and covered by the project's existing checks.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 58/100