Azure / Azure/MachineLearningNotebooks

PyTorch on Azure ML: problem accessing GPU with ACPT environment image

Đang mở
#1,917 12 bình luận 1 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Jupyter Notebook
Star
4.4k
Fork
2.6k
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

In Azure Machine Learning, I am trying to set up a **PyTorch 2.0.1 run** on a curated **ACPT image**:
`mcr.microsoft.com/azureml/curated/acpt-pytorch-2.0-cuda11.7:latest`
(based on [this](https://learn.microsoft.com/en-us/azure/machine-learning/how-to-azure-container-for-pytorch-environment?view=azureml-api-2) and [this](https://github.com/Azure/AzureML-Containers/tree/master/base/gpu) documentation)

I use the image above on a GPU cluster with size `STANDARD_NC6`and this conda environment yaml:
```
name: compass-environment-simple
channels:
- anaconda
- pytorch
- nvidia
- conda-forge
dependencies:
- python==3.10.9
- pip
- pytorch==2.0.1
- torchvision
- torchaudio
- pytorch-cuda=11.7
- cudatoolkit=11.7
- pip:
- pandas==2.0.1
- numpy==1.24.3
- azure-ai-ml==1.7.2
- azureml-mlflow==1.51.0
- mlflow==2.3.2
- tensorboard==2.13.0
- tensorboardX==2.6
- matplotlib==3.7.1
- matplotlib-inline==0.1.6
- mizani==0.9.1
- plotnine==0.12.1
- seaborn==0.12.2
- scikit-learn==1.2.2
- scipy==1.10.1
- numpy==1.24.3
- plotly==5.14.1
- statsmodels==0.14.0
- ipython==8.14.0
- tqdm==4.65.0
- umap-learn==0.5.3
- imbalanced-learn==0.10.1
- gensim==4.3.1
- nltk==3.8.1
- python-dotenv==1.0.0
- openpyxl==3.1.2
- xlwt==1.3.0
- xlrd==2.0.1
- XlsxWriter==3.1.0
```

But I keep getting a couple of error messages and I cannot access the GPUs when running the PyTorch Python scripts.

1) while building the image
in `20_image_build_log.txt` I get a warning that the NVIDIA driver can't be detected:
`WARNING: The NVIDIA Driver was not detected. GPU functionality will not be available.`
(see the attached [20_image_build_log.txt](https://github.com/Azure/MachineLearningNotebooks/files/11771108/20_image_build_log.txt) at line 3408)

2) when executing the Python script during a run
I get the following error message:
```
/bin/bash: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/libtinfo.so.6: no version information available (required by /bin/bash)
['/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python310.zip', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/lib-dynload', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/mpmath-1.2.1-py3.10.egg', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b']
Traceback (most recent call last):
File "/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd/training_script.py", line 22, in
import torch
File "/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/__init__.py", line 229, in
from torch._C import * # noqa: F403
ImportError: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so: undefined symbol: cudaGraphInstantiateWithFlags, version libcudart.so.11.0
```

What am I doing wrong?
Thanks for your help.

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.