Azure / Azure/MachineLearningNotebooks

Pandas dataframes with array column values are not correctly persisted as AzureML datasets

Ouverte
#1,587 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
ADO bug MLOps
Langage dominant
Jupyter Notebook
Étoiles
4.4k
Forks
2.6k
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

Pandas dataframes with arrays as column values seem to be incorrectly persisted. An example:

```python
test_df = pd.DataFrame({'x': [np.random.rand(1000) for _ in range(1000)]})
ds = Datastore.get_default(ws)
Dataset.Tabular.register_pandas_dataframe(test_df, ds, 'test_dataset')

test_df.head()
###
x
0 [0.5044850335733219, 0.6054305053424696, 0.669...
1 [0.41759815476145723, 0.266477750018155, 0.511...
2 [0.6777708610872593, 0.16925324567267985, 0.16...
3 [0.4268294269387616, 0.6540643485117185, 0.033...
4 [0.6560106490417036, 0.5804652379458484, 0.582...

Dataset.get_by_name(ws, 'test_dataset').to_pandas_dataframe().head()
###
x
0 ERROR
1 ERROR
2 ERROR
3 ERROR
4 ERROR
```

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.