A column-vector y was passed when a 1d array was expected (however, y is already a 1d array)
Nobody has claimed this yet.
- Dominant language
- Jupyter Notebook
- Stars
- 4.8k
- Forks
- 827
- PR merge metrics
- No merged PRs in 30d
Description
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier, RandomForestRegressor
from econml.dml import NonParamDML
X, y = make_classification(random_state=42)
T = y.copy()
y, T = y.ravel(), T.ravel()
causal_model = NonParamDML(model_y=RandomForestClassifier(),
model_t=RandomForestClassifier(),
model_final=RandomForestRegressor(),
discrete_outcome=True,
discrete_treatment=True,
cv=3)
causal_model.fit(y, T, X=X)
successfully fit while displaying the following warning message:
A column-vector y was passed when a 1d array was expected. Please change the shape of y to (n_samples,), for example using ravel().
Is there anything incorrect in the code or are there any ways to avoid the warning?
Thanks a lot for any feedback.
Best,
Yanis
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the provided NonParamDML example and tracing how its fit inputs reach the scikit-learn models. Identify which input produces the column-vector warning, then confirm the expected behavior with a focused regression test; done means the valid 1D inputs no longer emit the misleading warning.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, scikit-learn
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100