alteryx / alteryx/evalml

Spike: Preserve woodwork schema through dask

Abierto
#2,252 2 comentarios 1 reacción 1 asignado Reclamado por @asniyaz Ver en GitHub
refactor spike
Lenguaje dominante
Python
Estrellas
850
Forks
96
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

In #2243, we added `X_schema` and `y_schema` arguments to `train_pipeline` and `score_pipeline` because the dask engine would lose the schema once it pickles the dataframe and sends it to the workers.

The underlying problem is that `pickle` does not preserve the woodwork schema. This sounds like a general accessor thing and is probably not specific to woodwork.

```python
import pandas as pd
import woodwork as ww
import pickle
import pytest

df = pd.DataFrame({"a": [1, 2, 3], 'id': [0, 1, 2]}, index=[4, 5, 6])
df.ww.init()
assert df.ww.schema is not None

assert pickle.loads(pickle.dumps(df)).ww.schema is None
```

This issue tracks figuring out how to make dask serialize woodwork data structures. Once we come up with a solution, we can discuss if this should live in evalml or woodwork.

Some resources to check out:

- [Extending dask serialization](https://distributed.dask.org/en/latest/serialization.html#id3)
- [woodwork serializers](https://woodwork.alteryx.com/en/stable/generated/woodwork.serialize.typing_info_to_dict.html#woodwork.serialize.typing_info_to_dict)

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.