alteryx / alteryx/evalml

Standardize how we access unique target values for classification problems

Open
#3,112 2 comments 0 reactions 1 assignee Claimed by @asniyaz View on GitHub
refactor tech debt
Dominant language
Python
Stars
850
Forks
96
PR merge metrics
No merged PRs in 30d

Description

Right now, we access unique target values for classification problems in several ways:
1. `list(ww.init_series(np.unique(y)))` (`classification_pipeline.py`)
2. `unique_labels` (`confusion_matrix`)
3. LabelBinarizer / np.unique in `roc_curve` (slightly different than label encoding)

It could be helpful to standardize how we encode and decode targets pre and post fit time. This issue tracks finding places where we encode/decode and seeing how we could standardize this process.

Note that in some cases, we might encode/decode outside of the context of a pipeline (such as confusion_matrix), but it could still be helpful to consolidate our implementation to fewer methods if possible!

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.