Azure / Azure/MachineLearningNotebooks
OutputDatasetConfig.register_on_complete registers dataset if the step finish with error
- Lenguaje dominante
- Jupyter Notebook
- Estrellas
- 4.4k
- Forks
- 2.6k
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
While I was running a pipeline a step finished with **Error**:
AzureMLCompute job failed.
DiskFullError: Disk full while running job. Reduce amount of data accessed, or upgrade VM Sku.
As a result of this step I had defined an OutputDatasetConfig with the properties "as_upload" and "register_on_complete". What I was expecting was not to upload dataset neither register it because the step finished with error, so the output is not right, but the situation was that the dataset was upload an registered, and this implies that a tagged version of the dataset is corrupted.
I recommend not to register a dataset if the step finishes with an error that it's what I would expect from documentation.
Regards
---
#### Document Details
⚠ *Do not edit this section. It is required for docs.microsoft.com ➟ GitHub issue linking.*
* ID: 02631223-bb1d-f9de-2536-23d753c98508
* Version Independent ID: f524ca56-5419-b233-b67a-a1b3d10408e7
* Content: [azureml.data.output_dataset_config.OutputDatasetConfig class - Azure Machine Learning Python](https://docs.microsoft.com/en-us/python/api/azureml-core/azureml.data.output_dataset_config.outputdatasetconfig?view=azure-ml-py)
* Content Source: [AzureML-Docset/stable/docs-ref-autogen/azureml-core/azureml.data.output_dataset_config.OutputDatasetConfig.yml](https://github.com/MicrosoftDocs/MachineLearning-Python-pr/blob/live/AzureML-Docset/stable/docs-ref-autogen/azureml-core/azureml.data.output_dataset_config.OutputDatasetConfig.yml)
* Service: **machine-learning**
* Sub-service: **core**
* GitHub Login: @DebFro
* Microsoft Alias: **debfro**
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
Comienza con la documentación de OutputDatasetConfig y la fuente de AzureML-Docset enlazada; después, reproduce un paso de pipeline que falle al usar as_upload y register_on_complete. Traza cuándo se producen la carga y el registro después del fallo. Se considera terminado cuando los pasos fallidos no cargan ni registran datasets de salida dañados, con cobertura para el caso de error informado.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- azure, machine-learning, python
- Área
- cloud, machine-learning
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Estancado
- Claridad
- Bastante claro
- Aptitud para principiantes
- 35/100