envoyproxy / envoyproxy/gateway
Best Practices for Load Shedding in Envoy Gateway
- Dominant language
- Go
- Stars
- 3k
- Forks
- 864
- Avg merge
- 1d 22h
- Merged PRs (30d)
- 148
Description
Hello,
I'm looking for guidance on implementing load shedding in Envoy Gateway to protect backend services under high load.
In networking, mechanisms such as Random Early Detection (RED) and traffic policing proactively prevent congestion by dropping or limiting traffic before the network becomes saturated. I'm looking for the equivalent approach for HTTP/API traffic using Envoy Gateway.
Specifically, I'd like to know:
Does Envoy Gateway support proactive load shedding based on resource pressure (e.g., request latency, concurrency, queue depth, CPU utilization, or other overload signals)?
Is Envoy's adaptive concurrency filter currently supported and configurable through Envoy Gateway?
What is the recommended way to reject excess requests before backend services become overloaded?
Are there best practices for combining:
Local or global rate limiting
Circuit breakers
Adaptive concurrency
Overload Manager
Kubernetes HPA
If some of these capabilities are not yet exposed by Envoy Gateway, what is the recommended production approach today?
Our goal is not only to enforce rate limits, but to gracefully shed load when the platform approaches its safe operating capacity, ensuring that the system remains responsive instead of allowing latency to grow until services become unavailable.
If there are existing examples, documentation, or recommended configuration patterns for this use case, I would greatly appreciate being pointed to them.
Thank you!
Contributor guide
No contributing guide indexed for this repository
Research direction
No file, test, or entry point is identified in the issue. Start by reviewing Envoy Gateway's existing documentation and configuration support for rate limiting, circuit breakers, adaptive concurrency, Overload Manager, and Kubernetes HPA; done means documenting which capabilities are supported and the recommended production load-shedding approach.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- kubernetes
- Domain
- backend-api-design, networking
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100