containerd / containerd/nri

Documentation: Best practices for zero-downtime plugin updates

Open
#167 3 comments 3 reactions 0 assignees View on GitHub
documentation
Dominant language
Go
Stars
406
Forks
102
Avg merge
1d 10h
Merged PRs (30d)
8

Description

I have an NRI plugin running as a DaemonSet. Ideally, I would like to set `spec.updateStrategy.rollingUpdate.maxSurge` to a positive number, and `spec.updateStrategy.rollingUpdate.maxUnavailable` to zero to have "make before break" semantics where on each node the new pod starts up and becomes ready, before the old pod is terminated.

However, for my NRI plugin, it's not safe to have two instances running at the same time acting on the same container, since this would result in an action being performed twice on pod startup rather than once. So for now, I am fully terminating the old pod before starting the new pod to prevent overall. And when the new pod comes up, it has to "catch up" to process any events that were missed.

It would be helpful if we could document any known patterns for achieving this kind of zero-downtime update scenario for an NRI plugin, in a way that avoids duplicate processing.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.