aws / aws/eks-anywhere

Baremetal: The cluster status is not restored to the state before the failure, when upgrading nodes failed

Open
#5,254 0 comments 0 reactions 0 assignees View on GitHub
area/upgrades external
Dominant language
Go
Stars
2.1k
Forks
328
Avg merge
1d 4h
Merged PRs (30d)
9

Description

**What happened**:
I tried to upgrade the nodes of the eksa cluster, and failed.
The status of nodes and the phase of machines are not restored to the status before the failure, when upgrading nodes failed.

1. status of the nodes after the failure: Ready, SchedulingDisabled (uncordon command can't change the status)
2. phase of the machines: Deleting

**What you expected to happen**:
I expected the cluster status including node status and machine phase would restored to the status before failure

**How to reproduce it (as minimally and precisely as possible)**:
```bash
# 1. upgrade the nodes
eksctl anywhere upgrade cluster -f eksa-mgmt-cluster-upgrade.yaml --kubeconfig ~/.kube/config -v 9

# 2. upgrade failed
Error: failed to upgrade cluster: waiting for workload cluster control plane to be ready: executing wait: executing wait: error: timed out waiting for the condition on clusters/eksa-bm

# 3. check the status of nodes and machines
root@eksaadm:~/cluster/eksa-create-log# kubectl get nodes
NAME STATUS ROLES AGE VERSION
eksacp01 Ready control-plane 4h18m v1.24.6-eks-4360b32
eksacp02 Ready control-plane 4h10m v1.24.6-eks-4360b32
eksacp03 Ready,SchedulingDisabled control-plane 4h2m v1.24.6-eks-4360b32
eksadp01 Ready 4h8m v1.24.6-eks-4360b32
eksadp02 Ready 4h9m v1.24.6-eks-4360b32
root@eksaadm:~/cluster/eksa-create-log# kubectl get machines -A
NAMESPACE NAME CLUSTER NODENAME PROVIDERID PHASE AGE VERSION
eksa-system eksa-bm-55ss7 eksa-bm eksacp03 tinkerbell://eksa-system/eksacp03 Deleting 4h v1.24.10-eks-1-24-9
eksa-system eksa-bm-cqd8g eksa-bm eksacp02 tinkerbell://eksa-system/eksacp02 Running 4h v1.24.10-eks-1-24-9
eksa-system eksa-bm-dmhc6 eksa-bm eksacp01 tinkerbell://eksa-system/eksacp01 Running 4h v1.24.10-eks-1-24-9
eksa-system eksa-bm-md-0-5c5bc585ff-b9dz4 eksa-bm eksadp01 tinkerbell://eksa-system/eksadp01 Running 4h v1.24.10-eks-1-24-9
eksa-system eksa-bm-md-0-5c5bc585ff-tlghg eksa-bm eksadp02 tinkerbell://eksa-system/eksadp02 Running 4h v1.24.10-eks-1-24-9
```

**Anything else we need to know?**:
- error message when the upgrading the nodes failed
```bash
2023-03-15T11:36:38.190+0900 V9 docker {"stderr": "error: timed out waiting for the condition on clusters/eksa-bm\n"}
2023-03-15T11:36:38.190+0900 V5 Error happened during retry {"error": "executing wait: error: timed out waiting for the condition on clusters/eksa-bm\n", "retries": 1}
2023-03-15T11:36:38.191+0900 V5 Execution aborted by retry policy
2023-03-15T11:36:38.191+0900 V4 Task finished {"task_name": "upgrade-workload-cluster", "duration": "1h30m44.050965922s"}
2023-03-15T11:36:38.191+0900 V4 ----------------------------------
2023-03-15T11:36:38.191+0900 V4 Task start {"task_name": "collect-cluster-diagnostics"}
2023-03-15T11:36:38.191+0900 V0 collecting cluster diagnostics
2023-03-15T11:36:38.191+0900 V0 collecting management cluster diagnostics
2023-03-15T11:36:38.191+0900 V0 collecting workload cluster diagnostics
2023-03-15T11:36:38.213+0900 V3 bundle config written {"path": "eksa-bm/generated/eksa-bm-2023-03-15T11:36:38+09:00-bundle.yaml"}
2023-03-15T11:36:38.213+0900 V1 creating temporary namespace for diagnostic collector {"namespace": "eksa-diagnostics"}
2023-03-15T11:36:38.213+0900 V5 Retrier: {"timeout": "2562047h47m16.854775807s", "backoffFactor": null}
2023-03-15T11:36:38.214+0900 V6 Executing command {"cmd": "/snap/bin/docker exec -i eksa_1678842339032055239 kubectl create namespace eksa-diagnostics --kubeconfig eksa-bm/eksa-bm-eks-a-cluster.kubeconfig"}
2023-03-15T11:36:38.538+0900 V5 Retry execution successful {"retries": 1, "duration": "325.003901ms"}
2023-03-15T11:36:38.539+0900 V1 creating temporary ClusterRole and RoleBinding for diagnostic collector
2023-03-15T11:36:38.539+0900 V5 Retrier: {"timeout": "2562047h47m16.854775807s", "backoffFactor": null}
2023-03-15T11:36:38.539+0900 V6 Executing command {"cmd": "/snap/bin/docker exec -i eksa_1678842339032055239 kubectl apply -f - --kubeconfig eksa-bm/eksa-bm-eks-a-cluster.kubeconfig"}
2023-03-15T11:36:39.267+0900 V5 Retry execution successful {"retries": 1, "duration": "728.525448ms"}
2023-03-15T11:36:39.267+0900 V0 ⏳ Collecting support bundle from cluster, this can take a while {"cluster": "eksa-bm", "bundle": "eksa-bm/generated/eksa-bm-2023-03-15T11:36:38+09:00-bundle.yaml", "since": "2023-03-15T08:36:38.213+0900", "kubeconfig": "eksa-bm/eksa-bm-eks-a-cluster.kubeconfig"}
2023-03-15T11:36:39.268+0900 V6 Executing command {"cmd": "/snap/bin/docker exec -i eksa_1678842339032055239 support-bundle eksa-bm/generated/eksa-bm-2023-03-15T11:36:38+09:00-bundle.yaml --kubeconfig eksa-bm/eksa-bm-eks-a-cluster.kubeconfig --interactive=false --since-time 2023-03-15T08:36:38.21385285+09:00"}
2023-03-15T11:38:58.902+0900 V0 Support bundle archive created {"path": "support-bundle-2023-03-15T02_36_39.tar.gz"}
2023-03-15T11:38:58.902+0900 V0 Analyzing support bundle {"bundle": "eksa-bm/generated/eksa-bm-2023-03-15T11:36:38+09:00-bundle.yaml", "archive": "support-bundle-2023-03-15T02_36_39.tar.gz"}
2023-03-15T11:38:58.902+0900 V6 Executing command {"cmd": "/snap/bin/docker exec -i eksa_1678842339032055239 support-bundle analyze eksa-bm/generated/eksa-bm-2023-03-15T11:36:38+09:00-bundle.yaml --bundle support-bundle-2023-03-15T02_36_39.tar.gz --output json"}
2023-03-15T11:39:00.212+0900 V0 Analysis output generated {"path": "eksa-bm/generated/eksa-bm-2023-03-15T11:39:00+09:00-analysis.yaml"}
2023-03-15T11:39:00.212+0900 V1 cleaning up temporary roles for diagnostic collectors
2023-03-15T11:39:00.212+0900 V5 Retrier: {"timeout": "2562047h47m16.854775807s", "backoffFactor": null}
2023-03-15T11:39:00.212+0900 V6 Executing command {"cmd": "/snap/bin/docker exec -i eksa_1678842339032055239 kubectl delete -f - --kubeconfig eksa-bm/eksa-bm-eks-a-cluster.kubeconfig"}
2023-03-15T11:39:00.530+0900 V5 Retry execution successful {"retries": 1, "duration": "318.195134ms"}
2023-03-15T11:39:00.530+0900 V1 cleaning up temporary namespace for diagnostic collectors {"namespace": "eksa-diagnostics"}
2023-03-15T11:39:00.530+0900 V5 Retrier: {"timeout": "2562047h47m16.854775807s", "backoffFactor": null}
2023-03-15T11:39:00.530+0900 V6 Executing command {"cmd": "/snap/bin/docker exec -i eksa_1678842339032055239 kubectl delete namespace eksa-diagnostics --kubeconfig eksa-bm/eksa-bm-eks-a-cluster.kubeconfig"}
2023-03-15T11:39:07.080+0900 V5 Retry execution successful {"retries": 1, "duration": "6.549966059s"}
2023-03-15T11:39:07.080+0900 V4 Task finished {"task_name": "collect-cluster-diagnostics", "duration": "2m28.889499555s"}
2023-03-15T11:39:07.080+0900 V4 ----------------------------------
2023-03-15T11:39:07.080+0900 V4 Saving checkpoint {"file": "eksa-bm-checkpoint.yaml"}
2023-03-15T11:39:07.081+0900 V4 Tasks completed {"duration": "1h33m25.867481495s"}
2023-03-15T11:39:07.081+0900 V3 Cleaning up long running container {"name": "eksa_1678842339032055239"}
2023-03-15T11:39:07.082+0900 V6 Executing command {"cmd": "/snap/bin/docker rm -f -v eksa_1678842339032055239"}
Error: failed to upgrade cluster: waiting for workload cluster control plane to be ready: executing wait: executing wait: error: timed out waiting for the condition on clusters/eksa-bm
```

**Environment**:
- EKS Anywhere Release: 0.14.2
- EKS Distro Release: 1.24

Contributor guide

Open the contributing guide

Research direction

Start at the `eksctl anywhere upgrade cluster` entry point and inspect the failed-upgrade path, using `eksa-bm-checkpoint.yaml` and the reported node and Machine states as evidence. Reproduce the timeout on a bare-metal cluster and verify that node scheduling and Machine phase return to their pre-upgrade values after failure.

Written by the indexing model from the issue text.

Assessment

Tech stack
docker, go, kubernetes
Domain
devops, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.