alteryx / alteryx/evalml

Spike: Preserve woodwork schema through dask

Open
#2,252 2 comments 1 reaction 1 assignee Claimed by @asniyaz View on GitHub
refactor spike
Dominant language
Python
Stars
850
Forks
96
PR merge metrics
No merged PRs in 30d

Description

In #2243, we added `X_schema` and `y_schema` arguments to `train_pipeline` and `score_pipeline` because the dask engine would lose the schema once it pickles the dataframe and sends it to the workers.

The underlying problem is that `pickle` does not preserve the woodwork schema. This sounds like a general accessor thing and is probably not specific to woodwork.

```python
import pandas as pd
import woodwork as ww
import pickle
import pytest

df = pd.DataFrame({"a": [1, 2, 3], 'id': [0, 1, 2]}, index=[4, 5, 6])
df.ww.init()
assert df.ww.schema is not None

assert pickle.loads(pickle.dumps(df)).ww.schema is None
```

This issue tracks figuring out how to make dask serialize woodwork data structures. Once we come up with a solution, we can discuss if this should live in evalml or woodwork.

Some resources to check out:

- [Extending dask serialization](https://distributed.dask.org/en/latest/serialization.html#id3)
- [woodwork serializers](https://woodwork.alteryx.com/en/stable/generated/woodwork.serialize.typing_info_to_dict.html#woodwork.serialize.typing_info_to_dict)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.