alteryx / alteryx/evalml

Spike: Preserve woodwork schema through dask

Đang mở
#2,252 2 bình luận 1 reaction 1 người được giao Được @asniyaz nhận Xem trên GitHub
refactor spike
Ngôn ngữ chính
Python
Star
850
Fork
96
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

In #2243, we added `X_schema` and `y_schema` arguments to `train_pipeline` and `score_pipeline` because the dask engine would lose the schema once it pickles the dataframe and sends it to the workers.

The underlying problem is that `pickle` does not preserve the woodwork schema. This sounds like a general accessor thing and is probably not specific to woodwork.

```python
import pandas as pd
import woodwork as ww
import pickle
import pytest

df = pd.DataFrame({"a": [1, 2, 3], 'id': [0, 1, 2]}, index=[4, 5, 6])
df.ww.init()
assert df.ww.schema is not None

assert pickle.loads(pickle.dumps(df)).ww.schema is None
```

This issue tracks figuring out how to make dask serialize woodwork data structures. Once we come up with a solution, we can discuss if this should live in evalml or woodwork.

Some resources to check out:

- [Extending dask serialization](https://distributed.dask.org/en/latest/serialization.html#id3)
- [woodwork serializers](https://woodwork.alteryx.com/en/stable/generated/woodwork.serialize.typing_info_to_dict.html#woodwork.serialize.typing_info_to_dict)

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.