Azure / Azure/MachineLearningNotebooks
PyTorch on Azure ML: problem accessing GPU with ACPT environment image
- Dominant language
- Jupyter Notebook
- Stars
- 4.4k
- Forks
- 2.6k
- PR merge metrics
- No merged PRs in 30d
Description
In Azure Machine Learning, I am trying to set up a **PyTorch 2.0.1 run** on a curated **ACPT image**:
`mcr.microsoft.com/azureml/curated/acpt-pytorch-2.0-cuda11.7:latest`
(based on [this](https://learn.microsoft.com/en-us/azure/machine-learning/how-to-azure-container-for-pytorch-environment?view=azureml-api-2) and [this](https://github.com/Azure/AzureML-Containers/tree/master/base/gpu) documentation)
I use the image above on a GPU cluster with size `STANDARD_NC6`and this conda environment yaml:
```
name: compass-environment-simple
channels:
- anaconda
- pytorch
- nvidia
- conda-forge
dependencies:
- python==3.10.9
- pip
- pytorch==2.0.1
- torchvision
- torchaudio
- pytorch-cuda=11.7
- cudatoolkit=11.7
- pip:
- pandas==2.0.1
- numpy==1.24.3
- azure-ai-ml==1.7.2
- azureml-mlflow==1.51.0
- mlflow==2.3.2
- tensorboard==2.13.0
- tensorboardX==2.6
- matplotlib==3.7.1
- matplotlib-inline==0.1.6
- mizani==0.9.1
- plotnine==0.12.1
- seaborn==0.12.2
- scikit-learn==1.2.2
- scipy==1.10.1
- numpy==1.24.3
- plotly==5.14.1
- statsmodels==0.14.0
- ipython==8.14.0
- tqdm==4.65.0
- umap-learn==0.5.3
- imbalanced-learn==0.10.1
- gensim==4.3.1
- nltk==3.8.1
- python-dotenv==1.0.0
- openpyxl==3.1.2
- xlwt==1.3.0
- xlrd==2.0.1
- XlsxWriter==3.1.0
```
But I keep getting a couple of error messages and I cannot access the GPUs when running the PyTorch Python scripts.
1) while building the image
in `20_image_build_log.txt` I get a warning that the NVIDIA driver can't be detected:
`WARNING: The NVIDIA Driver was not detected. GPU functionality will not be available.`
(see the attached [20_image_build_log.txt](https://github.com/Azure/MachineLearningNotebooks/files/11771108/20_image_build_log.txt) at line 3408)
2) when executing the Python script during a run
I get the following error message:
```
/bin/bash: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/libtinfo.so.6: no version information available (required by /bin/bash)
['/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python310.zip', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/lib-dynload', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages', '/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/mpmath-1.2.1-py3.10.egg', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe', '/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b']
Traceback (most recent call last):
File "/mnt/azureml/cr/j/7db89bab2dab4c479669caa5a4af805b/exe/wd/training_script.py", line 22, in
import torch
File "/azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/__init__.py", line 229, in
from torch._C import * # noqa: F403
ImportError: /azureml-envs/azureml_8536d86a92a3687bb89e1cff176aedf4/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so: undefined symbol: cudaGraphInstantiateWithFlags, version libcudart.so.11.0
```
What am I doing wrong?
Thanks for your help.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.