4paradigm / 4paradigm/k8s-vgpu-scheduler

core dump when request 2 or more gpus with Tesla T4

Open
#24 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
595
Forks
100
PR merge metrics
No merged PRs in 30d

Description

### 1. Issue or feature description
It's ok when request 1 gpu in yaml. But when request more than 1, the output of nvidia-smi is below:
![image](https://user-images.githubusercontent.com/22440678/176585153-3a43f4fa-c8be-4693-b3bd-38cc125711f2.png)
The output of nvidia-smi in host machine is ok.

In another machine with GeForce RTX 2070 SUPER ,it's all right when request 2 gpus.
![image](https://user-images.githubusercontent.com/22440678/176585865-ca820791-a21c-4b5c-b8b9-5e626ca18205.png)
but when I run application locally , it abort due to :
```
[4pdvGPU ERROR (pid:697 thread=140106827071488 context.c:189)]: cuCtxGetDevice Not Found. tid=140106827071488 ctx=0x239601906000:0x23960041a000
home/limengxuan/work/libcuda_override/src/cuda/context.c:189: cuCtxGetDevice: Assertion `0' failed.
```

### 2. Steps to reproduce the issue
ubuntu1~20.04 + microk8s + Tesla T4 GPU + 510driver
### 3. Information to [attach](https://help.github.com/articles/file-attachments-on-issues-and-pull-requests/) (optional if deemed irrelevant)

Common error checking:
- [ ] The output of `nvidia-smi -a` on your host
![image](https://user-images.githubusercontent.com/22440678/176586282-c3a0c48d-1a8d-4faf-85ca-1bff87e4a78b.png)
- [ ] Your docker configuration file (e.g: `/etc/docker/daemon.json`)
-{
"default-runtime": "nvidia",
"runtimes": {
"nvidia": {
"path": "nvidia-container-runtime",
"runtimeArgs": []
}
}
}

Additional information that might help better understand your environment and reproduce the bug:
- [ ] Any relevant kernel output lines from `dmesg`
```
nvidia-smi[2260220]: segfault at 0 ip 00007fde46d051ce sp 00007ffe1ae4c9e8 error 4 in libc-2.31.so[7fde46b9d000+178000]
[89993.700532] Code: fd d7 c9 0f bc d1 c5 fe 7f 27 c5 fe 7f 6f 20 c5 fe 7f 77 40 c5 fe 7f 7f 60 49 83 c0 1f 49 29 d0 48 8d 7c 17 61 e9 c2 04 00 00 fe 6f 1e c5 fe 6f 56 20 c5 fd 74 cb c5 fd d7 d1 49 83 f8 21 0f
[90182.697502] nvidia-smi[2265941]: segfault at 0 ip 00007f241971c1ce sp 00007fffff703d08 error 4 in libc-2.31.so[7f24195b4000+178000]
[90182.697509] Code: fd d7 c9 0f bc d1 c5 fe 7f 27 c5 fe 7f 6f 20 c5 fe 7f 77 40 c5 fe 7f 7f 60 49 83 c0 1f 49 29 d0 48 8d 7c 17 61 e9 c2 04 00 00 fe 6f 1e c5 fe 6f 56 20 c5 fd 74 cb c5 fd d7 d1 49 83 f8 21 0f

```

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue involves a core dump when requesting multiple Tesla T4 GPUs via the k8s-vgpu-scheduler. Start by examining the device plugin code handling multi-GPU allocation, particularly around CUDA context management. Check logs and the provided dmesg output for segmentation faults in nvidia-smi. Reproduce the environment with Ubuntu 20.04, microk8s, Tesla T4, and driver 510. Look at the nvidia-container-runtime configuration and how the scheduler maps multiple GPUs to pods.

Written by the indexing model from the issue text.

Assessment

Tech stack
docker, kubernetes
Domain
cloud, devops, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.