scikit-learn / scikit-learn/scikit-learn

RFC SLEP006: verbose vs non-verbose declaration in meta-estimator

Open
#23,928 16 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

API RFC
Dominant language
Python
Stars
67.3k
Forks
27.4k
Avg merge
1d 15h
Merged PRs (30d)
58

Description

As the proposal and the implementation of meta-estimator routing (SLEP006) stands, if the user wants to use sample_weight, they need to be quite verbose in how they declare the estimators. Taking AdaBoostClassifier as an example, and imagining if AdaBoostClassifier would use the sub-estimator's score method, the user would have to write:

est = (
    AdaBoostClassifier(LogisticRegression().set_fit_request(sample_weight=True)
    .set_score_request(sample_weight=True))
    .fit(X, y, sample_weight=sw)
)

which is quite more verbose than the current code users need to write:

est = AdaBoostClassifier(LogisticRegression()).fit(X, y, sample_weight=sw)

There have been concerns about making users write quite verbose code in cases where the current pattern seems quite reasonable.

Without changing everything related to SLEP006, there are three paths we can take:

Option 1: Helper function

We can introduce helper functions to make the above code simpler. For instance, a weighted function could request sample_weight on all methods which accept sample_weight for a given estimator. Then the above code would look like:

est = AdaBoostClassifier(weighted(LogisticRegression())).fit(X, y, sample_weight=sw)

and if the sub-estimator is a pipeline:

est = AdaBoostClassifier(
    make_pipeline(weighted(StandardScaler()), weighted(LogisticRegression())))
).fit(X, y, sample_weight=sw)

Implementing weighted for a Pipeline (or other meta-estimators) would be tricky since set_fit_request is only available for consumers and not non-consumer routers; therefore the user needs to repeat the weighted call for all sub-estimators.

Option 2: Different meta-estimators

Have two classes of meta-estimators (or routers to be specific).

In this scenario, we divide meta-estimators into two classes, simple and complex. Simple routers are the ones which simply forward **kwargs to sub-estimators, and by default the assume sub-estimators have requested those metadata. This simplifies the users' code and makes the existing code for simple meta-estimators to keep working, but it raises a few issues.

First is that there will be two classes of meta-estimators, and the user would need to know which estimator is of which class. It's also not clear what we should do if the user explicitly sets request values for metadata (we can probably respect those if present).

Another issue is that if a meta-estimator changes behavior, it needs to become a complex meta-estimator if we want to keep backward compatibility for it. This doesn't seem like a good pattern.

Option 3: Keep as is

Do nothing, things are as is.

I'm in favor or option 1 because:

  • with the helper function the user code doesn't look too verbose
  • using metadata is not a beginner kinda thing and therefore this API is not hampering beginners' experience with the library
  • it keeps consistency among meta-estimators/consumers

xref: https://github.com/scikit-learn/scikit-learn/pull/22986#pullrequestreview-1005144559

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading the SLEP006 proposal and the linked pull request review. Compare the three options in the issue and confirm the intended API direction with maintainers before locating implementation work. Done requires an agreed approach for reducing or retaining declaration verbosity, plus the corresponding implementation and tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
api, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.