awslabs / awslabs/eks-perf-tests

KIT Guest Clusters ETCD pods unable to become ready when using Karpenter v0.11.0+

Open
#241 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

0.2 bug
Dominant language
Shell
Stars
72
Forks
50
PR merge metrics
No merged PRs in 30d

Description

When using Karpenter version v0.12.1, using a Provisioner that has a small ttlSecondsAfterEmpty can result in removing the node that an ETCD replica is scheduled to. The associated PVC for the ETCD pod then never binds to the volume.

The initial thought is since Karpenter does not pre-bind to pods anymore after v0.11.0, this may introduce some undesired behavior for the EBS Volumes on the ETCD instances.

The nodes for the PVCs for ETCD 1 and 2 were never deleted. Even though Karpenter brought up a replacement node for ETCD 0, the pod and PVC never binded

NAME                                 STATUS    VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   AGE
etcd-data-kit-guest-cluster-etcd-0   Pending                                                                        kit-gp3        23m
etcd-data-kit-guest-cluster-etcd-1   Bound     pvc-f993bc10-327d-4414-a3b2-acc95d956ab0   40Gi       RWO            kit-gp3        23m
etcd-data-kit-guest-cluster-etcd-2   Bound     pvc-f781eb15-4bfb-4395-b2c1-52b53c54a63a   40Gi       RWO            kit-gp3        23m

➜  k get pv
NAME                                       CAPACITY   ACCESS MODES   RECLAIM POLICY   STATUS   CLAIM                                             STORAGECLASS   REASON   AGE
pvc-f781eb15-4bfb-4395-b2c1-52b53c54a63a   40Gi       RWO            Delete           Bound    tekton-tests/etcd-data-kit-guest-cluster-etcd-2   kit-gp3                 22m
pvc-f993bc10-327d-4414-a3b2-acc95d956ab0   40Gi       RWO            Delete           Bound    tekton-tests/etcd-data-kit-guest-cluster-etcd-1   kit-gp3                 22m

This is the error message when describing the ETCD pod.

  Warning  FailedScheduling  20m                 default-scheduler  running PreBind plugin "VolumeBinding": binding volumes: failed to get node "ip-192-168-86-66.us-west-2.compute.internal": node "ip-192-168-86-66.us-west-2.compute.internal" not found
  Warning  FailedScheduling  16m (x2 over 18m)   default-scheduler  (combined from similar events): 0/5 nodes are available: 1 node(s) didn't find available persistent volumes to bind, 2 node(s) didn't have free ports for the requested pod ports, 2 node(s) didn't match Pod's node affinity/selector.
  Warning  FailedScheduling  15m (x6 over 19m)   default-scheduler  0/5 nodes are available: 1 Insufficient cpu, 1 Insufficient memory, 1 Too many pods, 2 node(s) didn't have free ports for the requested pod ports, 2 node(s) didn't match Pod's node affinity/selector.
  Warning  FailedScheduling  15m (x6 over 19m)   default-scheduler  0/5 nodes are available: 1 node(s) had taint {node.kubernetes.io/not-ready: }, that the pod didn't tolerate, 2 node(s) didn't have free ports for the requested pod ports, 2 node(s) didn't match Pod's node affinity/selector.
  Warning  FailedScheduling  11m (x5 over 19m)   default-scheduler  0/4 nodes are available: 2 node(s) didn't have free ports for the requested pod ports, 2 node(s) didn't match Pod's node affinity/selector.
  Warning  FailedScheduling  97s (x20 over 18m)  default-scheduler  0/5 nodes are available: 1 node(s) didn't find available persistent volumes to bind, 2 node(s) didn't have free ports for the requested pod ports, 2 node(s) didn't match Pod's node affinity/selector.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the linked Karpenter v0.11.0 upgrade guide, then inspect the ETCD pod and PVC scheduling events reported here in a Kubernetes cluster using Karpenter v0.12.1 and a short ttlSecondsAfterEmpty. Determine whether node removal leaves the PVC tied to a missing node; done means the ETCD replica PVC binds and the pod becomes Ready after replacement-node provisioning.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, kubernetes
Domain
cloud, databases, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.