py-why / py-why/EconML

Shapley Values from CausalForestDML

Open
#751 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
4.8k
Forks
827
PR merge metrics
No merged PRs in 30d

Description

I am working with the Customer Segmentation notebook found here.

I wanted to look at the Shapley values for X and W's in the data set and get the constant marginal effects. The code below shows how I ran the model and produced the charts. This leads to my question given the model DAG assumptions mentioned in this thread. This leads to my questions which are related to each other.

  1. If the second chart shows the effect of the W's on Y (demand), why wouldn't the chart also include the effect of income (X) and price (T)?
  2. Given the question above, is there a way to recover the Shapley values that account for for the effects of X,T and W's? I have looked at the Shapley values in the dict's that are produced. They look just fine but each dictionary has a different set of base values which indicates to me these are being run (as indicated in the code) as completely separate models. Is there a proper or recommended way to recover the Shaps from a single model? My goal is to use the Shapley values to say "this W accounts for this much movement away from the base value, while X accounts for this much...".

import shap
from econml.dml import CausalForestDML
est = CausalForestDML(model_y=GradientBoostingRegressor(), model_t=GradientBoostingRegressor())

est.tune(log_Y, log_T, X=X, W=W)
est.fit(log_Y, log_T, X=X, W=W)
shap_values_w = est.shap_values(W)
shap_values_x = est.shap_values(X)
shap.plots.beeswarm(shap_values_x['demand']['price'])
shap.plots.beeswarm(shap_values_w['demand']['price'])

xy

wy

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the Customer Segmentation notebook and the CausalForestDML calls shown in this issue, especially est.shap_values(X) and est.shap_values(W); review the related discussion in issue 744. Determine whether the existing API can represent combined X, T, and W contributions from one model, and define the expected result before considering implementation or documentation.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook, python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.