GoogleCloudPlatform / GoogleCloudPlatform/esp-v2

Healthz returns incorrect status with health_check_grpc_backend active

Open
#751 4 comments 0 reactions 1 assignee Claimed by @qiwzhang View on GitHub
Dominant language
Go
Stars
307
Forks
185
Avg merge
14h 45m
Merged PRs (30d)
6

Description

Hello
I am using espv2 (2.39.0 ) active health checking to track backend status.
Parameters I use are:

--healthz=healthz
--health_check_grpc_backend
--health_check_grpc_backend_interval=5s

My expectation is:
when ESPv2 has started and backend service has not started yet an esp endpoint /healthz should fail.
However a response received is 200 OK with body { "code": 200, "message": "" }

In espv2 container log I periodically see

In espv2 container log I periodically see lines
"D1112 20:22:43.232 24 D1112 20:44:18.625 27 external/envoy/source/common/http/codec_client.cc:57] [27][client][C23] connecting
D1112 20:44:18.625 27 external/envoy/source/common/network/connection_impl.cc:924] [27][connection][C23] connecting to 192.168.65.2:21411
D1112 20:44:18.625 27 external/envoy/source/common/network/connection_impl.cc:943] [27][connection][C23] connection in progress
D1112 20:44:20.657 27 external/envoy/source/common/network/connection_impl.cc:695] [27][connection][C23] delayed connect error: 111
D1112 20:44:20.657 27 external/envoy/source/common/network/connection_impl.cc:250] [27][connection][C23] closing socket: 0
D1112 20:44:20.657 27 external/envoy/source/common/http/codec_client.cc:108] [27][client][C23] disconnect. resetting 1 pending requests
D1112 20:44:20.657 27 external/envoy/source/common/http/codec_client.cc:140] [27][client][C23] request reset
D1112 20:44:20.657 27 external/envoy/source/common/upstream/health_checker_impl.cc:787] [27][hc][C23] connection/stream error health_flags=healthy

Is that an intended behavior?

I use ingress load balancer which is monitoring /healthz endpoint to check for status. A pod is either added or removed from a loadbalancer based on the status. With the current setup (on rolling update) on rollout a failing pod is added to a balancer causing request timeouts, because backend is not ready yet but esp tells that backend pod is ready. After a couple of seconds a pod is removed from a load balancer after esp does a couple of checks to a backend and change status from healthy to unhealthy.

Thanks!

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.