py-why / py-why/EconML

CausalForestDML does not accept my inputs with NaN during training, even with a featurizer that deals with NaN

Open
#989 7 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
4.8k
Forks
827
PR merge metrics
No merged PRs in 30d

Description

Here is a code to see the problem:

import numpy as np
import pandas as pd
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from econml.dml import CausalForestDML
from sklearn.ensemble import RandomForestRegressor, RandomForestClassifier
from sklearn.model_selection import train_test_split

# Step 1: Create synthetic dataset
np.random.seed(0)
n = 300
X = pd.DataFrame({
    'age': np.random.normal(40, 10, size=n),
    'income': np.random.normal(50000, 15000, size=n),
    'education_years': np.random.normal(16, 2, size=n)
})

# Introduce missing values in 'income'
missing_idx = np.random.choice(n, size=30, replace=False)
X.loc[missing_idx, 'income'] = np.nan

# Binary treatment and outcome
T = np.random.binomial(1, 0.5, size=n)
Y = 1000 + 500*T + 20*X['age'] + np.random.normal(0, 1000, size=n)

# Step 2: Featurizer to handle missing data
featurizer = Pipeline([
    ('imputer', SimpleImputer(strategy='mean')),
    ('scaler', StandardScaler())
])

# Step 3: CausalForestDML model
est = CausalForestDML(
    model_t=RandomForestClassifier(),
    model_y=RandomForestRegressor(),
    featurizer=featurizer,
    discrete_treatment=True,
    random_state=0
)

# Step 4: Fit the model
est.fit(Y, T, X=X)

Unfortunately this throws an error:

ValueError: Input contains NaN.

The issue appears to be that during fit(), CausalForestDML runs the following:

X, = check_input_arrays(X, force_all_finite='allow-nan' if 'X' in self._gen_allowed_missing_vars() else True)

Which isn't really right because it should rather check the featurized X which is what is ultimately used for fitting anyways, not X itself.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the issue with the provided CausalForestDML example, then inspect fit() and the check_input_arrays call shown in the report. Determine how validation interacts with the featurizer, and verify that fitting with NaN values succeeds when the featurizer handles them without weakening unrelated input checks.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, scikit-learn
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.