hashicorp / hashicorp/consul

Wrong Instance Count because of Cluster Restart

Open
#17,013 7 comments 0 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
30.1k
Forks
4.6k
Avg merge
2d 6h
Merged PRs (30d)
43

Description

#### Overview of the Issue
When a Azure Cluster is restarted, the previous pods running inside the mesh are lost and new pods are created. Post restart Consul isn't picking up that the previous pods are deleted/not present. Hence, it is still trying to route to the pods resulting in the following error when I try to consume the service be UI/API.

![image](https://user-images.githubusercontent.com/80001447/232262847-a378318a-6906-4e93-95d9-22aa2908a03b.png)

Further more:
These are the active pods:
![image](https://user-images.githubusercontent.com/80001447/232262916-4e253b0c-2fee-4f5d-ba52-f1f7b7fcf059.png)

Whereas, Consul UI shows this,
![image](https://user-images.githubusercontent.com/80001447/232262931-500c9d88-dba2-4a37-9760-c3e2004bb893.png)

Notice that there is only one frontend pod is running in AKS whereas UI shows two instances of the service

---

#### Reproduction Steps

Steps to reproduce this issue, eg:

1. Create a service mesh within Azure Kubernetes Service.
1. Stop and Start the AKS.
1. Notice the previous pod info still exists in Consul UI whereas in reality it doesn't exist in AKS.

-->

### Consul info for both Client and Server

Client info
/ $ consul info
agent:
check_monitors = 0
check_ttls = 0
checks = 0
services = 0
build:
prerelease =
revision = 7c04b6a0
version = 1.15.1
version_metadata =
consul:
acl = disabled
bootstrap = true
known_datacenters = 1
leader = true
leader_addr = 10.244.0.12:8300
server = true
raft:
applied_index = 2213
commit_index = 2213
fsm_pending = 0
last_contact = 0
last_log_index = 2213
last_log_term = 4
last_snapshot_index = 0
last_snapshot_term = 0
latest_configuration = [{Suffrage:Voter ID:b9744a41-cccd-861f-eca2-f3b18496e5b4 Address:10.244.0.12:8300}]
latest_configuration_index = 0
num_peers = 0
protocol_version = 3
protocol_version_max = 3
protocol_version_min = 0
snapshot_version_max = 1
snapshot_version_min = 0
state = Leader
term = 4
runtime:
arch = amd64
cpu_count = 4
goroutines = 259
max_procs = 4
os = linux
version = go1.20.1
serf_lan:
coordinate_resets = 0
encrypted = false
event_queue = 1
event_time = 4
failed = 0
health_score = 0
intent_queue = 1
left = 0
member_time = 4
members = 1
query_queue = 0
query_time = 1
serf_wan:
coordinate_resets = 0
encrypted = false
event_queue = 0
event_time = 1
failed = 0
health_score = 0
intent_queue = 0
left = 0
member_time = 1
members = 1
query_queue = 0
query_time = 1

### Operating system and Environment details
Kubernetes Version: 1.24.10
Cloud Provider: Azure
Environment: Azure Kubernetes Service

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.