py-why / py-why/EconML

score_nuisances with discrete treatment returns incorrect score

Open Beginner friendly
#1,006 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
4.8k
Forks
827
PR merge metrics
No merged PRs in 30d

Description

When using score_nuisances with a discrete treatment, the function does not return the correct score.

The issue comes from the inverse_onehot function in econml/utilities.py. Currently, when it receives as input a DataFrame generated by pandas.get_dummies(), it incorrectly decodes the treatment.

For example, in case of binary treatments, labels originally coded as 0 and 1 are shifted and end up being decoded as 1 and 2, due to the following implementation:

def inverse_onehot(T):
    """
    Given a one-hot encoding of a value, return a vector reversing the encoding to get numeric treatment indices.

    Note that we assume that the first column has been removed from the input.
    """
    assert ndim(T) == 2
    # note that by default OneHotEncoder returns float64s, so need to convert to int
    return (T @ np.arange(1, T.shape[1] + 1)).astype(int)

This logic introduces an off-by-one error when decoding treatments.

Expected behavior

The function should return zero-based indices, ensuring that discrete treatments (e.g. 0/1) remain consistent after decoding. A corrected implementation would have the following code:

def inverse_onehot(T):
    assert econml.utilities.ndim(T) == 2

    indices = (
        np.arange(0, T.shape[1])
        if isinstance(T, pd.DataFrame)
        else np.arange(1, T.shape[1] + 1)
    )

    return (T @ indices).astype(int)

This change guarantees that score_nuisances computes the correct score for discrete treatments.

Contributed by @Cantal00p, @f5ilverio

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in econml/utilities.py at inverse_onehot and trace how score_nuisances passes discrete-treatment DataFrames to it. Check the pandas.get_dummies() and non-DataFrame paths, then verify that zero-based treatment labels are preserved and score_nuisances returns the correct score for binary treatments.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
76/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.