Azure / Azure/MachineLearningNotebooks

AutoML experiment: get model and metrics for any algorithm (not only for the best one)

Abierto
#1,885 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Jupyter Notebook
Estrellas
4.4k
Forks
2.6k
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

## What I'm trying to do

For an **AutoML Forecasting** experiment, I'd like to compare the performance of the **best model** with the performance of **another model** from the same experiment.

For an AutoML run, I understand how to get the best performing model and its metrics like this:
```
# ...initialize MLFlow client...
mlflow_parent_run = mlflow_client.get_run('upbeat_square_abs3942')
best_child_run_id = mlflow_parent_run.data.tags["automl_best_child_run_id"]
best_run = mlflow_client.get_run(best_child_run_id)
best_run.data.metrics
# etc...
```

But how can I fetch the job for _any_ model based on the _algorithm name_?
Something like:
```
# pseudocode:
mlflow_client.get_automl_run_by_algorithm('XGBoostRegressor')
```

## So far, I managed to figure out the following:

1. list of algorithms used in the AutoML experiment
```
mlflow_parent_run.data.tags['pipeline_id_000']
# '__AutoML_Naive__;__AutoML_SeasonalNaive__;__AutoML_Average__;__AutoML_SeasonalAverage__;__AutoML_Ensemble__'
```
However, this list seems to be in an arbitrary order and I struggle to get the corresponding job names for the algorithms.

2. "internal" job names for the child runs

The child runs seem to have different names than the names shown in Azure ML Studio.
They are named for instance `upbeat_square_abs3942_2` - i.e. the name of the parent run `upbeat_square_abs3942` followed by **underscore plus a number** (`_2`in this example).

But Azure ML Studio displays names like (no `upbeat_square_abs3942_2` to be found):
image
So this code works:
```
child_run = mlflow_client.get_run('upbeat_square_abs3942_2')
```
but using a name shown in the screenshot above throws an exception, e.g.
```
child_run = mlflow_client.get_run('green_floor_0ln3tlpv')
```

### Question

How can I obtain the model and metrics for any algorithm used in the experiment?

Thanks!

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Comienza con el uso del cliente de MLflow mostrado en el issue, incluidos get_run, la etiqueta automl_best_child_run_id y pipeline_id_000. Rastrea cómo se identifican las ejecuciones secundarias de AutoML y cómo se relacionan los nombres de los algoritmos con sus nombres de ejecución internos. Se considera completado cuando se documente o habilite la recuperación de un modelo y sus métricas para cualquier algoritmo del experimento, no solo para la mejor ejecución.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
azure, jupyter-notebook, python
Área
api, machine-learning
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
35/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.