aws / aws/containers-roadmap

[EKS] [request]: AWS Health Event Automatic Remediation for the managed nodegroups

Open
#2,005 0 comments 0 reactions 0 assignees View on GitHub
EKS EKS Managed Nodes Proposed
Dominant language
Shell
Stars
5.4k
Forks
334
PR merge metrics
No merged PRs in 30d

Description

### Community Note

* Please vote on this issue by adding a 👍 [reaction](https://blog.github.com/2016-03-10-add-reactions-to-pull-requests-issues-and-comments/) to the original issue to help the community and maintainers prioritize this request
* Please do not leave "+1" or "me too" comments, they generate extra noise for issue followers and do not help prioritize the request
* If you are interested in working on this issue or have submitted a pull request, please leave a comment

**Tell us about your request**
What do you want us to build?
Time to time the AWS deprecates the EC2, its internal hardware or failures in the AZs or retire the EC2, these can be part of a managed nodegroup of an EKS cluster. These events comes to the customer as part of AWS Health dashboard notifications. Customers has to manage such nodes manually, we want to do it automatically. The MNGs should take-care of such nodes programmatically as anyway this will be enforced on the customer after the given date.

Additionally, the team can build it in such a way that it should be an opt-in functionality that we would need to explicitly configure when creating the ManagedNodeGroup resources. If we tell you it's safe to do it, and it turns out not to be, that's customer's fault. Kubernetes as a product offers enough features to allow us to make it safe.

**Which service(s) is this request for?**
EKS

**Tell us about the problem you're trying to solve. What are you trying to do, and why is it hard?**
If it's not solved than customers have to do it manually so that their applications running on the certain nodes doesn't get impacted. Many times customers with bigger or multiple EKS clusters may have this issue where they have to use a dedicated resource to look into this node replacement, which can be easily be solved using automated process. Also, since it's "MANAGED" by AWS, why can't an EC2 service can talk to EKS service and do it anyway?

**Are you currently working around this issue?**
Manually

**Additional context**
Anything else we should know?

**Attachments**
If you think you might have additional information that you'd like to include via an attachment, please do - we'll take a look. (Remember to remove any personally-identifiable information.)

Contributor guide

Open the contributing guide

Research direction

Start by reviewing how AWS Health dashboard notifications identify deprecated, failing, or retiring EC2 instances in EKS managed nodegroups. The request does not name files, tests, or an implementation entry point. Done would require a defined, safe, explicitly opt-in automatic remediation flow for affected nodes.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, kubernetes
Domain
cloud, infrastructure
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.