dotnet / dotnet/machinelearning-modelbuilder

Column purposes should have Sampling Key

Open
#1,873 1 comment 1 reaction 1 assignee Claimed by @LittleLittleCloud View on GitHub
Feature request Priority:1 Reported by: Customer Stale
Dominant language
Dockerfile
Stars
285
Forks
66
PR merge metrics
No merged PRs in 30d

Description

**Is your feature request related to a problem? Please describe.**
At the moment, Advanced Data Options dialog box only supports 3 column purposes: Feature / Label / Ignore.

Some datasets overfit easily when validation/test data is chosen randomly. These might include datasets which include items about the same entity (for example, blood results of the same patient at different times or credit card fraud transactions dataset where one person has commited multiple fraud) or time based series for related items (for example, will it rain tomorrow in a specific location and dataset only include one country). Running this kinds of experiments with Model Builder will often lead to overfitted models, which blocks usage of Model builder for these purposes.

**Describe the solution you'd like**
Add "Sampling Key" purpose to Advanced Data Options.

**Additional context**
This has been been suggested earlier by justinormont in 2019 https://github.com/dotnet/machinelearning-modelbuilder/issues/75#issuecomment-506642972 but I did not notice discussion about it not being suitable for ML Builder, so maybe it could be now reconsidered?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.