Standardize how we access unique target values for classification problems
- Dominant language
- Python
- Stars
- 850
- Forks
- 96
- PR merge metrics
- No merged PRs in 30d
Description
Right now, we access unique target values for classification problems in several ways:
1. `list(ww.init_series(np.unique(y)))` (`classification_pipeline.py`)
2. `unique_labels` (`confusion_matrix`)
3. LabelBinarizer / np.unique in `roc_curve` (slightly different than label encoding)
It could be helpful to standardize how we encode and decode targets pre and post fit time. This issue tracks finding places where we encode/decode and seeing how we could standardize this process.
Note that in some cases, we might encode/decode outside of the context of a pipeline (such as confusion_matrix), but it could still be helpful to consolidate our implementation to fewer methods if possible!
Contributor guide
Assessment
This issue has not been assessed yet.