Azure / Azure/MachineLearningNotebooks

ParallelRunStep with CSV input introduces duplicate rows

Open
#1,703 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
4.4k
Forks
2.6k
PR merge metrics
No merged PRs in 30d

Description

Using AzureML, ParallelRunStep with a CSV file input, I run Shapley explanations (SHAP package) in parallel on a cluster. I see the actual mini batches run are greater than expected, leading to extra rows in the output file.

The cluster is a 10 node Standard_D64_v3 created 2/28/2022 (I have destroyed and recreated the cluster and seen the same issue).

The metric Total Mini Batches is present by default. I also log the batch size in my script, and in this way I get one metric per mini batch.
Here are my observations:

* The input CSV has no duplicate row IDs
* The file produced by ParallelRunStep has duplicate IDs
* The run completes successfully
* The number of the batch size entries is greater than total mini batches expected. The difference between these 2 is usually the exact number of extra rows.

I work around this by forcing uniqueness in the output file; however this is not ideal for large datasets.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.