alibaba / alibaba/x-deeplearning
通过k8s 来启动单个xdl 实例报错
- Dominant language
- PureBasic
- Stars
- 4.3k
- Forks
- 1k
- PR merge metrics
- No merged PRs in 30d
Description
```ruby
Error: failed to start container "kml-dtmachine-520": Error response from daemon:
OCI runtime create failed: container_linux.go:348:
starting container process caused "process_linux.go:402:
container init caused \"process_linux.go:385:
running prestart hook 0 caused \\\"error running hook:
exit status 1, stdout: , stderr: exec command:
[/usr/bin/nvidia-container-cli --load-kmods configure --ldconfig=@/sbin/ldconfig --device=GPU-73b2a28a-071f-dfe9-b5a4-8a648de3fdc4 --utility --pid=26740 /media/disk1/docker/overlay/c287354d7e8e0d349d8aa3eb92f7cecae7145b1ec19fbc42ec57c2ba9d8b7830/merged]
\\\\nnvidia-container-cli: mount error: file creation failed:
/media/disk1/docker/overlay/c287354d7e8e0d349d8aa3eb92f7cecae7145b1ec19fbc42ec57c2ba9d8b7830/merged/usr/bin/nvidia-smi:
file exists\\\\n\\\"\"": unknown
```
我在物理上手动起没问题,通过k8s 调度起有问题
```
k8s v1.11
driver 396.44
default runtime nvidia-docker
```
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the reported OCI runtime error and compare the manual launch with the Kubernetes path. Check the Kubernetes 1.11 setup, nvidia-docker default runtime, and driver 396.44 configuration; done means a single xdl instance starts through Kubernetes without the nvidia-smi file-exists error.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- docker, kubernetes
- Domain
- infrastructure
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100