Azure / Azure/MachineLearningNotebooks

Use Hyperdrive to optimize pipeline hyperparameters

Aperta
#1,526 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
ADO bug Training
Lingua principale
Jupyter Notebook
Stelle
4.4k
Fork
2.6k
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

I would like to use Hyperdrive to optimize a full pipeline. That is, I would like to optimize hyperparameters on different steps jointly. I raised the issue here https://github.com/MicrosoftDocs/azure-docs/issues/77227 but I was suggested to open it here too.

For example, I have a pipeline defined as:
```
[prepare_data]
|
v
[extract_features]
|
v
[train_model]
```
I can use Hyperdriver to tune the hyperparameters of my ML model in the `train_model` step based upon some metrics, say validation loss. What I would like to do is to tune the hyperparameters in the `train_model` step together with the hyperparameters in the pre-processing steps (e.g., `extract_features`). For example, I would like to find the best _sequence length_ in `extract_features` that can improve the loss in the model training.

[`HyperDriveConfig`](https://docs.microsoft.com/en-us/python/api/azureml-train-core/azureml.train.hyperdrive.hyperdriveconfig?view=azure-ml-py) does accept an argument `pipeline`, which seems to be exactly what I am looking for. Unfortunately, I cannot find much information on how to use this parameter.

I tried to submit a Hyperdriver run as:
```python
hd_config = HyperDriveConfig(
hyperparameter_sampling=...,
policy=...,
primary_metric_name=...,
primary_metric_goal=...,
max_total_runs=...,
max_duration_minutes=...,
max_concurrent_runs=...,
pipeline=pipeline,
)

exp = Experiment(workspace=ws, name="test")
hd_run = exp.submit(hd_config)
```
where `pipeline` is one of my published pipelines in the workspace that accepts PipelineParameters to tune. However, I get the error:
```
Exception has occurred: AttributeError
'PublishedPipeline' object has no attribute 'graph'
```
How could I proceed?

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

Inizia con l’argomento pipeline di HyperDriveConfig e il percorso di invio di PublishedPipeline descritti nell’issue, quindi riproduci l’AttributeError con la configurazione Python mostrata. Controlla l’issue collegata relativa alla documentazione di Azure e gli esempi di pipeline esistenti per individuare il tipo di pipeline previsto e il flusso dei parametri supportato. Il lavoro è completato quando viene documentato o implementato un modo supportato per ottimizzare congiuntamente i parametri di preprocessing e training e l’errore segnalato viene risolto.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
azure, machine-learning, python
Ambito
cloud, machine-learning
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.