envoyproxy / envoyproxy/envoy

Cluster healthcheck uses wrong interval due to metric filtering

Open
#8,771 6 comments 0 reactions 0 assignees View on GitHub
bug help wanted
Dominant language
C++
Stars
28.9k
Forks
5.6k
Avg merge
1d 20h
Merged PRs (30d)
428

Description

Given an **Envoy v1.11.1** node with [cluster](https://www.envoyproxy.io/docs/envoy/latest/configuration/upstream/cluster_manager/cluster_stats) metrics excluded via `StatsMatcher` and cluster health checking enabled, when sending traffic to the cluster healthcheck is always run with `no_traffic_interval` 60s default interval.

The cause is the same as with https://github.com/envoyproxy/envoy/issues/8473 and https://github.com/envoyproxy/envoy/issues/8630, metrics are being used for application functionality.

In this case [upstream_cx_total](https://github.com/envoyproxy/envoy/blob/v1.11.1/source/common/upstream/health_checker_base_impl.cc#L86) is the metric on which the dependency exists. Therefore the workaround is to whitelist it:

```
stats_config:
stats_matcher:
inclusion_list:
patterns:
- suffix: upstream_cx_total
```

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.