py-why / py-why/EconML

MemoryError in econml.dml

Open
#331 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Jupyter Notebook
Stars
4.8k
Forks
827
PR merge metrics
No merged PRs in 30d

Description

I have a dataset with about 6.3 million observations (rows), 10 treatments and 4 effect modifiers. I am using econml.dml.LinearDML module like this:

est = LinearDMLCateEstimator(model_t =MultiOutputRegressor(CatBoostRegressor()), model_y = CatBoostRegressor()) est.fit(Y, T, X, W, inference='statsmodels')
And it throws a MemoryError (see below). It seems that in the inference step it creates a big matrix who crashes the memory. If I instead use econml.dml.KernelDML, the memory consumption is even larger.

I'm wondering if there's any more memory efficient way of implementation which would avoid the error.

File "/local/home/haohupku/run.py", line 162, in _get_coef est.fit(Y, T, X, W, inference='statsmodels') File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/utilities.py", line 1212, in m return to_wrap(*args, **kwargs) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/dml.py", line 608, in fit inference=inference) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/utilities.py", line 1212, in m return to_wrap(*args, **kwargs) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/dml.py", line 488, in fit inference=inference) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/utilities.py", line 1212, in m return to_wrap(*args, **kwargs) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/_rlearner.py", line 322, in fit inference=inference) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/utilities.py", line 1212, in m return to_wrap(*args, **kwargs) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/cate_estimator.py", line 104, in call m(self, Y, T, *args, **kwargs) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/_ortho_learner.py", line 551, in fit sample_var=self._subinds_check_none(sample_var, fitted_inds)) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/_ortho_learner.py", line 621, in _fit_final sample_var=sample_var)) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/_rlearner.py", line 107, in score effects = self._model_final.predict(X).reshape((-1, Y_res.shape[1], T_res.shape[1])) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/dml.py", line 217, in predict prediction = self._model.predict(self._combine(None if X is None else X2, T, fitting=False)) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/dml.py", line 161, in _combine return cross_product(F, T) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/utilities.py", line 293, in cross_product return _apply(cross, XS) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/utilities.py", line 233, in _apply result = op(*XS) File "/home/haohupku/projects/pkgs/miniconda/envs/py36/lib/python3.6/site-packages/econml/utilities.py", line 292, in cross return reshape(reduce(np.multiply, XS), (n, -1)) MemoryError: Unable to allocate 19.7 GiB for an array with shape (53004520, 10, 5) and data type float64

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with econml/dml.py, especially _combine and predict, then inspect econml/utilities.py around cross_product and the _rlearner.py scoring path shown in the traceback. Reproduce the 6.3-million-row fit and determine how inference constructs the oversized array; done means the failing workflow completes without the reported allocation, with regression coverage for the memory-sensitive path.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.