[EKS] [request]: Reset Managed Node Group status
- Dominant language
- Shell
- Stars
- 5.4k
- Forks
- 334
- PR merge metrics
- No merged PRs in 30d
Description
### Community Note
* Please vote on this issue by adding a 👍 [reaction](https://blog.github.com/2016-03-10-add-reactions-to-pull-requests-issues-and-comments/) to the original issue to help the community and maintainers prioritize this request
* Please do not leave "+1" or "me too" comments, they generate extra noise for issue followers and do not help prioritize the request
* If you are interested in working on this issue or have submitted a pull request, please leave a comment
**Tell us about your request**
We would like to have an alternative way to make an EKS Managed Node Group active/healthy again after a scale out failure.
**Which service(s) is this request for?**
EKS
**Tell us about the problem you're trying to solve. What are you trying to do, and why is it hard?**
Right now, when an EKS Managed Node Group does not succeed to scale out (due to capacity unavailability), its status becomes degraded. The only way to reset its status back to active/healthy is to successfully scale out. Even if we scale this EKS Managed Node Group to zero, its status remains degraded. So in case that AWS runs out of capacity for the specific instance type/az of this EKS Managed Node Group, it would remain in degraded status forever.
We would like to have an alternative way to reset its status back to active/healthy status.
The problem is that we need to make some exceptions to our tools/monitoring e.g. for EKS Managed Node Group scaled to zero but in degraded status.
**Are you currently working around this issue?**
We implement some very custom exceptions to our tools/monitoring.
Contributor guide
Research direction
This is a public roadmap request and names no repository files, tests, or implementation entry points. Start by reviewing the EKS Managed Node Group status behavior and the reported scale-out failure scenario; done would require an AWS-supported way to restore an affected group's active or healthy status without a successful scale-out.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, kubernetes
- Domain
- cloud, infrastructure
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100