envoyproxy / envoyproxy/envoy

Different drain time for POST drain_listeners vs LDS replacement/update drain time

Open
#34,500 10 comments 0 reactions 0 assignees View on GitHub
area/drain area/hot_restart enhancement no stalebot
Dominant language
C++
Stars
28.9k
Forks
5.6k
Avg merge
1d 22h
Merged PRs (30d)
430

Description

# Current situation

Envoy offers a few existing options around listener draining during hot-restart and process shutdown.

* `--drain-strategy` allows a choice between `immediate` or `gradual` (i.e drain-time-s) draining for hot-restarts.
* `--drain-time-s` governs how long graceful draining (sending H2 `GOAWAY` or H1 `Connection: Close`) will take if triggered.
* `POST /drain_listeners?graceful&skip_exit&inboundonly` allows an operator to tell Envoy to begin draining listeners, but still accept new connections

# Gap

A problem emerges where we handle tens of millions of persistent connections and have competing goals for hot-restart draining length vs envoy shutdown or /drain_listeners?graceful draining length.

For hot restart, where we update listener configs across cross thousands of Envoy nodes simultaneously, we want a long >1+ hours draining of old listeners to prevent a storm of websocket re-handshakes.

For envoy shutdown (via `POST /drain_listeners?graceful`), which happens 1 node at a time, we want a much steeper/shorter graceful draining slope of ~5-10 minutes. (relevant to Auto scaling EC2 shutdown time)

In both cases, we want envoy and the listeners to continue accepting connections and applying H2 `GOAWAY` or H1 `Connection: Close` to transactions. We don't want to immediately "close" listeners and reject connections/requests.

# Solutions

We're thinking of three possible solutions:
1. Add an additional query string argument to `POST /drain_listeners?graceful` which allows the operator to specify a shorter drain time. E.g `POST /drain_listeners?graceful&drain-time-s=300`
2. Add a `--drain-time-s-triggered` for `POST /drain_listeners?graceful` which allows the operator to specify a shorter drain time.
3. Allow a different choice of `immediate` vs `gradual` `drain-strategy` for hot-restarts vs calling `/drain_listeners` (this is probably acceptable, but less preferable)

Relevant docs:
https://www.envoyproxy.io/docs/envoy/latest/operations/cli#cmdoption-drain-strategy
https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/operations/draining
https://www.envoyproxy.io/docs/envoy/latest/operations/admin#operations-admin-interface-drain

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.