Kaggle / Kaggle/docker-python

Pytorch CUDA P100 GPU Incompatibility

Aberta
#1,546 4 comentários 2 reações 0 responsáveis Ver no GitHub
bug help wanted
Linguagem predominante
Python
Estrelas
2.7k
Forks
1k
Merge médio
7d 14h
PRs com merge (30d)
2

Descrição

## 🐛 Bug

The model fails to train on GPU due to a CUDA capability mismatch. The installed version of PyTorch requires CUDA capability sm_70 or higher, but the available GPU (Tesla P100-PCIE-16GB) only supports sm_60. Here is traceback:

```
/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:435: UserWarning:
Found GPU0 Tesla P100-PCIE-16GB which is of cuda capability 6.0.
Minimum and Maximum cuda capability supported by this version of PyTorch is
(7.0) - (12.0)

queued_call()
/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:435: UserWarning:
Please install PyTorch with a following CUDA
configurations: 12.6 following instructions at
https://pytorch.org/get-started/locally/

queued_call()
/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:435: UserWarning:
Tesla P100-PCIE-16GB with CUDA capability sm_60 is not compatible with the current PyTorch installation.
The current PyTorch install supports CUDA capabilities sm_70 sm_75 sm_80 sm_86 sm_90 sm_100 sm_120.
If you want to use the Tesla P100-PCIE-16GB GPU with PyTorch, please check the instructions at https://pytorch.org/get-started/locally/

queued_call()
```

### To Reproduce

Use P100 GPU on Kaggle, move Pytorch neural network model to GPU and start training loop.

### Expected behavior

Successful forward training pass without errors.

### Additional context

This was the error given:

```
AcceleratorError: CUDA error: no kernel image is available for execution on the device
Search for `cudaErrorNoKernelImageForDevice' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.
```

Code worked fine for me a few weeks ago, but I'm guessing a change in the Docker environment broke something.

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Direção de pesquisa

Reproduza o loop de treinamento em uma Kaggle Tesla P100 e inspecione as versões do PyTorch e do CUDA no ambiente Docker. Compare a configuração instalada com o recurso sm_60 compatível e as orientações do PyTorch para CUDA 12.6; considera-se concluído quando a imagem oferecer suporte a uma passagem de treinamento forward bem-sucedida na P100 ou documentar claramente a incompatibilidade e o ambiente necessário.

Escrita pelo modelo de indexação a partir do texto da issue.

Avaliação

Stack de tecnologia
docker, python, pytorch
Domínio
infrastructure, machine-learning
Tipo de issue
Bug
Dificuldade
4/5
Tempo estimado
3-5 dias
Status de atividade
Pouca atividade
Clareza
Razoavelmente clara
Facilidade para iniciantes
42/100

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.