Azure / Azure/MachineLearningNotebooks

Unable to load Hugging face transformers pytorch model on ML Notebook

オープン
#1,397 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る
ADO AML Compute Instance bug notebook Workspace Management
主要言語
Jupyter Notebook
スター
4.4k
フォーク
2.6k
PR マージ指標
30日以内にマージされた PR はありません

説明

I tried to test out MachineLearning Notebooks with my Google Colab run.
Created gpu compute (or cpu small) and download files locally from git
Using same hugging face transformers version as I use on colab, cannot load model on Azure ML Notebook.

Any chance to support Hugging Face transformers models in ML Notebooks or am I missing a point?

Torch is by default 1.6 so I upgrade that one too

**Azure Notebook: Python 3.6.9 :: Anaconda, Inc. **
```
import transformers
transformers.__version__
'4.4.2'
```

```
import torch
torch.__version__
##GPU T80 instance, also tried CPU instance too
'1.8.0'
```
```

#!sudo apt-get install git-lfs
!git clone https://huggingface.co/gorkemgoknar/gpt2-turkish-writer

#tokenizer loading is fine
from transformers import GPT2Tokenizer, AutoModelForCausalLM
import torch
tokenizer = GPT2Tokenizer.from_pretrained("content/chatbot/model_tr_writer/gpt2-turkish-writer",local_files_only=True)

model = AutoModelForCausalLM.from_pretrained("content/chatbot/model_tr_writer/gpt2-turkish-writer",local_files_only=True)

---------------------------------------------------------------------------
RuntimeError Traceback (most recent call last)
/anaconda/envs/azureml_py36/lib/python3.6/site-packages/transformers/modeling_utils.py in from_pretrained(cls, pretrained_model_name_or_path, *model_args, **kwargs)
1061 try:
-> 1062 state_dict = torch.load(resolved_archive_file, map_location="cpu")
1063 except Exception:

/anaconda/envs/azureml_py36/lib/python3.6/site-packages/torch/serialization.py in load(f, map_location, pickle_module, **pickle_load_args)
584 orig_position = opened_file.tell()
--> 585 with _open_zipfile_reader(opened_file) as opened_zipfile:
586 if _is_torchscript_zip(opened_zipfile):

/anaconda/envs/azureml_py36/lib/python3.6/site-packages/torch/serialization.py in __init__(self, name_or_buffer)
241 def __init__(self, name_or_buffer) -> None:
--> 242 super(_open_zipfile_reader, self).__init__(torch._C.PyTorchFileReader(name_or_buffer))
243

RuntimeError: [enforce fail at inline_container.cc:145] . PytorchStreamReader failed reading zip archive: failed finding central directory

During handling of the above exception, another exception occurred:

OSError Traceback (most recent call last)
in
----> 1 model = AutoModelForCausalLM.from_pretrained("content/chatbot/model_tr_writer/gpt2-turkish-writer",local_files_only=True)
2
3
4 # Get sequence length max of 1024
5 tokenizer.model_max_length=1024

/anaconda/envs/azureml_py36/lib/python3.6/site-packages/transformers/models/auto/modeling_auto.py in from_pretrained(cls, pretrained_model_name_or_path, *model_args, **kwargs)
1111 if type(config) in MODEL_FOR_CAUSAL_LM_MAPPING.keys():
1112 return MODEL_FOR_CAUSAL_LM_MAPPING[type(config)].from_pretrained(
-> 1113 pretrained_model_name_or_path, *model_args, config=config, **kwargs
1114 )
1115 raise ValueError(

/anaconda/envs/azureml_py36/lib/python3.6/site-packages/transformers/modeling_utils.py in from_pretrained(cls, pretrained_model_name_or_path, *model_args, **kwargs)
1063 except Exception:
1064 raise OSError(
-> 1065 f"Unable to load weights from pytorch checkpoint file for '{pretrained_model_name_or_path}' "
1066 f"at '{resolved_archive_file}'"
1067 "If you tried to load a PyTorch model from a TF 2.0 checkpoint, please set from_tf=True. "

OSError: Unable to load weights from pytorch checkpoint file for 'content/chatbot/model_tr_writer/gpt2-turkish-writer' at 'content/chatbot/model_tr_writer/gpt2-turkish-writer/pytorch_model.bin'If you tried to load a PyTorch model from a TF 2.0 checkpoint, please set from_tf=True.

```

**Google Colab - Python 3.7.10**
```
import torch
torch.__version__
"1.8.0+cu101"
```
```

import transformers
transformers.__version__
"4.4.2"

import torch
torch.__version__
"1.8.0+cu101"
```

Only differences seem torch version (yet I can run it on CPU on my Macos) and Python version 3.6.9 vs 3.7.10

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

まず、Issue に示されているローカルの gpt2-turkish-writer モデルパスを使用し、Python 3.6.9、transformers 4.4.2、torch 1.8.0 の Azure ML Notebook で失敗を再現します。ダウンロードした pytorch_model.bin と環境を Google Colab と比較し、その後、モデルの読み込みが安定して機能するかどうかを判断します。モデルが読み込まれるか、非互換性と必要なサポートが明確に文書化されれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
azure, huggingface, jupyter-notebook, python, pytorch
領域
machine-learning
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。