NVIDIA / NVIDIA/gpu-operator

containerd restart from nvidia-container-toolkit causes other daemonsets to get stuck

Open
#991 9 comments 4 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement needs-triage question
Dominant language
Go
Stars
2.9k
Forks
552
Avg merge
2d 4h
Merged PRs (30d)
90

Description

Original context and jounrnalctl logs here: https://github.com/containerd/containerd/issues/10437

As we know by default nvidia-container-toolkit sends a SIGHUP to containerd for the patched containerd config to take effect. Unfortunately the way gpu-operator schedules Daemonsets all at once, we have noticed our gpu discovery and nvidia device plugin pods get forever stuck in pending. This is primarily due to config-manager-init container getting stuck in Created and never transitioning to Running state due to containerd restart.

Timeline of race condition:

  • nvidia-container-toolkit and nvidia-device-plugin schedules
  • nvidia-device-plugin waits on toolkit-ready file validation via init container
  • Patches the config to update nvidia runtime
  • Sends SIGHUP and writes toolkit-ready file
  • config-manager-init container from nvidia-device-plugin pod enters Created state
  • containerd restarts
  • config-manager-init forever stuck in Created, hence device plugin never gets to start

Today the only way for us to recover is to manually delete the stuck daemonset pods.

While I understand at the core this is containerd issue but this has become so troublesome we are looking for entrypoint and node label hacks. We are willing to take a solution that allows us to modify the entrypoint configmaps of daemonsets managed by ClusterPolicy.

I think something similar was discovered here but different effect
https://github.com/NVIDIA/gpu-operator/commit/963b8dc87ed54632a7345c1fcfe842f4b7449565
and was fixed with a sleep

P.S. I am aware container-toolkit has an option to not restart containerd, but we need a restart for correct toolkit injection behavior

cc: @klueska

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the linked containerd issue and journalctl logs, then trace how ClusterPolicy-managed DaemonSets schedule config-manager-init and the GPU discovery or device plugin pods. Compare the SIGHUP timing with the transition from Created to Running. Done means defining and implementing a supported way to prevent or recover from pods stuck after the containerd restart.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes
Domain
devops, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.