Azure / Azure/MachineLearningNotebooks

Pipeline parameters used with DataPath and DataPathComputeBinding to specify side inputs of Parallel pipeline

未关闭
#1,801 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Jupyter Notebook
星标
4.4k
派生
2.6k
PR 合并指标
30 天内没有已合并 PR

描述

[Enter feedback here]
I'm following [this](https://docs.microsoft.com/en-us/python/api/azureml-pipeline-core/azureml.pipeline.core.graph.pipelineparameter?view=azure-ml-py) example to create a PipelineParameters for my Parallel pipeline

```
from azureml.core.datastore import Datastore
from azureml.data.datapath import DataPath, DataPathComputeBinding
from azureml.pipeline.steps import PythonScriptStep
from azureml.pipeline.core import PipelineParameter

datastore = Datastore(workspace=workspace, name="workspaceblobstore")
datapath = DataPath(datastore=datastore, path_on_datastore='input_data')
data_path_pipeline_param = (PipelineParameter(name="input_data", default_value=datapath),
DataPathComputeBinding(mode='mount'))

train_step = PythonScriptStep(script_name="train.py",
arguments=["--input", data_path_pipeline_param],
inputs=[data_path_pipeline_param],
compute_target=compute_target,
source_directory=project_folder)

```

This is my code to create the pipeline with the parameters

```
path = DataPath(datastore=default_store, path_on_datastore='path')
input_param= (PipelineParameter(name="param_name", default_value=path), DataPathComputeBinding(mode='mount'))

parallel_run_config = ParallelRunConfig(
source_directory=script_dir,
entry_script='script.py', # the user script to run against each input
partition_keys=['key'],
error_threshold=50,
output_action='append_row',
environment=environment,
compute_target=compute_target,
node_count=2,
run_invocation_timeout=1200
)

parallel_run_step = ParallelRunStep(
name='test-batch-inference',
inputs=[partition_input],
side_inputs=[input1, input2, input_param],
output=output_dir,
parallel_run_config=parallel_run_config,
arguments=['--input_param', input_param],
allow_reuse=False
)
```
And it raised this error:

```
Exception: Step input must be of any type: (, , , , , , ), found

```

I'm using azureml-core==1.40.0.post2, azureml-pipeline==1.40.0
It's seems like the sample code is not supported with these version? Before trying this datapath as pipeline parameter, I tried int type input and its just work fine

---
#### Document Details

⚠ *Do not edit this section. It is required for docs.microsoft.com ➟ GitHub issue linking.*

* ID: 8e3ec7f7-25c2-8f63-331c-2eb62ffb73c7
* Version Independent ID: 4e31dffb-12fd-85d9-a1a2-aa038017d075
* Content: [azureml.pipeline.core.graph.PipelineParameter class - Azure Machine Learning Python](https://docs.microsoft.com/en-us/python/api/azureml-pipeline-core/azureml.pipeline.core.graph.pipelineparameter?view=azure-ml-py)
* Content Source: [AzureML-Docset/stable/docs-ref-autogen/azureml-pipeline-core/azureml.pipeline.core.graph.PipelineParameter.yml](https://github.com/MicrosoftDocs/MachineLearning-Python-pr/blob/live/AzureML-Docset/stable/docs-ref-autogen/azureml-pipeline-core/azureml.pipeline.core.graph.PipelineParameter.yml)
* Service: **machine-learning**
* Sub-service: **core**
* GitHub Login: @DebFro
* Microsoft Alias: **debfro**

贡献指南

这个仓库没有索引到贡献指南

调研方向

从 PipelineParameter API 页面和链接的 AzureML-Docset 源文件开始,然后使用 azureml-core==1.40.0.post2 和 azureml-pipeline==1.40.0 重现该示例。将文档中 DataPath/DataPathComputeBinding 的用法与报告的 ParallelRunStep 错误进行比较;当文档或受支持版本指南准确反映该行为时,即视为完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
azure, python
领域
documentation, machine-learning
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。