aws / aws/containers-roadmap

[EKS] [bug]: PodDisruptionBudget not respected during node group version update

Open
#2,623 1 comment 4 reactions 0 assignees View on GitHub
EKS EKS Managed Nodes Proposed
Dominant language
Shell
Stars
5.4k
Forks
334
PR merge metrics
No merged PRs in 30d

Description

### Community Note

* Please vote on this issue by adding a 👍 [reaction](https://blog.github.com/2016-03-10-add-reactions-to-pull-requests-issues-and-comments/) to the original issue to help the community and maintainers prioritize this request
* Please do not leave "+1" or "me too" comments, they generate extra noise for issue followers and do not help prioritize the request
* If you are interested in working on this issue or have submitted a pull request, please leave a comment

**Tell us about your request**
During a recent node group version update (initiated at 22:07 in the below screenshots), we observed an issue with the NGINX Ingress Controller pods not coming up on the new nodes. The reason for this is out of scope for this issue, but the PodDisruptionBudgets were not doing their thing during this incident, even though the Update button clearly indicates they should:

Image

Please see our metrics data that clearly shows the healthy instances dropping from two to one around 22:10 at the start of the upgrade. Also Disruptions Allowed goes to 0 at that moment, at which EKS should no longer evict any of the pods / remove any more nodes.

![Image](https://github.com/user-attachments/assets/25f3b5cc-a8fb-42d8-bb71-d09b04dc9e79)

**Available nodes during the process**

Image

**NGINX Ingress pods not available for the duration of the incident**

Image

As you can see (more information available on request) the nodes get scaled up and down as expected, but nowhere the PDB is taken into consideration.

**Upgrade process shows no errors**

Image

**Which service(s) is this request for?**
EKS

Contributor guide

Open the contributing guide

Research direction

No repository files, tests, or code entry points are identified. Start by reviewing the reported EKS node group version update, PodDisruptionBudget metrics, and attached incident evidence; done would require a confirmed explanation and a defined change that prevents further evictions once disruptions allowed reaches zero.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, kubernetes
Domain
cloud, infrastructure
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.