kubernetes / kubernetes/website

Document that pods hanging in terminating if DRA driver is truly gone is WAI (re: 129402 discussion)

Open
#51,012 18 comments 0 reactions 1 assignee Assigned to @lauralorenz View on GitHub
lifecycle/frozen needs-triage wg/device-management
Dominant language
HTML
Stars
5.4k
Forks
15.7k
Avg merge
4d 18h
Merged PRs (30d)
204

Description

If the DRA driver is well and truly gone, despite all the retry and reconciliation loops, a pod will be stuck in Terminating for as long as its NodeUnprepareResources call has not been fulfilled without error, which is (currently) impossible without a kubelet connection to the driver.

This is also true for networking plugins (as discussed in https://github.com/kubernetes/kubernetes/issues/129402#issuecomment-2578651690), and volumes/CSI drivers (https://github.com/kubernetes/kubernetes/issues/129402#issuecomment-2579086354) which have external services that handle the cleanup asynchronously, and sometimes untracked, by the pod phasing. Device Plugins don't have this issue (though they are at risk of leaving stuff lying around -- per https://github.com/kubernetes/kubernetes/issues/129402#issuecomment-2653348951).

This issue is to document this behavior as it pertains to DRA and describe how it is WAI and what the mediation steps available to a cluster administrator are.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.