Azure / Azure/MachineLearningNotebooks

Parallel run example fail when executing with deserialization error

オープン
#1,961 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Jupyter Notebook
スター
4.4k
フォーク
2.6k
PR マージ指標
30日以内にマージされた PR はありません

説明

I am trying to run a python script on multiple data inputs parallel, for this i am using the batch processing examples provided in this repo. Specifically I am looking at this one: https://github.com/Azure/MachineLearningNotebooks/blob/master/how-to-use-azureml/machine-learning-pipelines/parallel-run/file-dataset-partition-per-folder.ipynb

Running the exact same script initially generates a lot of value error saying "This pipeline didn't have the RawDeserializer policy; can't deserialize" (screenshot 1) and finally raises the Image build failed error (screenshot 2). I cannot find anything at the logfile mentioned in the error code. I have tried using the same environment from my azureml vm but that also didn't help.

Can someone kindly help?

![image](https://github.com/Azure/MachineLearningNotebooks/assets/64039635/14fae036-83a1-4e18-87b2-a1c8e42cbbc3)

![Screenshot 2024-04-10 at 17 58 37](https://github.com/Azure/MachineLearningNotebooks/assets/64039635/38644a70-fa37-4c12-93f1-141361b148ac)

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

Issue にリンクされている file-dataset-partition-per-folder.ipynb の例から始め、同じバッチ処理スクリプトを複数の入力に対して実行します。RawDeserializer および image-build の失敗を再現した後、その例がこれらのエラーなしで正常に完了し、利用可能なログが得られることを確認します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
azure, jupyter-notebook, python
領域
cloud, machine-learning
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。