Azure / Azure/azureml-examples

Using ParallelRunStep output as an input to another step

Open
#961 3 comments 0 reactions 0 assignees View on GitHub
example request pipeline
Dominant language
Jupyter Notebook
Stars
2k
Forks
1.7k
Avg merge
18h 18m
Merged PRs (30d)
2

Description

## Using ParallelRunStep output as an input to another step

This is specific to Python SDK.

I'm attempting to use ParallelRunStep output as an input to another step which I haven't been able to see an example of anywhere. My use case is simple, I wish to save the output of the pipeline with some additional transforms as a csv.

The closest example I've been able to find is here: https://docs.microsoft.com/en-us/azure/machine-learning/tutorial-pipeline-batch-scoring-classification#download-and-review-output, but it has a bug, the delimiter for the "parallel_run_step.txt" is a whitespace not a colon.

Here is code that I managed to get working eventually in my transform.py that's run after ParallelRunStep

transform.py

```python
import pandas as pd
import os
import argparse
from azureml.core import Run

parser = argparse.ArgumentParser(description="Transform")
parser.add_argument('--output_path', dest="output_path", required=True)

args, _ = parser.parse_known_args()

run = Run.get_context()
input_dir = run.input_datasets["input_data"]

input_data_path = os.path.join(input_dir,"parallel_run_step.txt")

input_df = pd.read_csv(input_data_path, delimiter=" ", header=None)
# Transform
transformed_df.to_csv(os.path.join(args.output_path,"processed_data.csv"))
```

PythonScriptStep

```python
transform_step = PythonScriptStep(
source_directory=src_dir,
name="transform",
script_name="transform.py",
compute_target=compute_target,
runconfig=aml_run_config,
inputs=[parallel_step_output.as_input('input_data') ],
arguments=["--output_path", saved_output],
```

Would be great to provide a similar example in the documentation and fix the issue with the wrong delimiter in the already provided examples

Contributor guide

Open the contributing guide

Research direction

Review the linked batch-scoring classification tutorial and the ParallelRunStep-to-PythonScriptStep snippets in this issue. Verify the delimiter used for parallel_run_step.txt, then add a documentation example showing how to pass ParallelRunStep output to a following Python step and save transformed data as CSV.

Written by the indexing model from the issue text.

Assessment

Tech stack
azure, python
Domain
documentation, machine-learning
Issue type
Documentation
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.