envoyproxy / envoyproxy/gateway
Support for Backend priorities beyond active/passive levels
- Dominant language
- Go
- Stars
- 3k
- Forks
- 864
- Avg merge
- 2d 2h
- Merged PRs (30d)
- 140
Description
*Description*:
Currently, Envoy Gateway supports two levels of backend priorities: active and passive. While this is sufficient for basic use cases, it limits more advanced traffic management scenarios for GenAI traffic where multiple levels of backend priorities are required for multiple regions, model providers, on-demand and provisioned throughput with different context length.
For multi-region deployment, we want to prioritize backends in following orders to achieve higher SLO target.
- us-east-1
- us-east-2
- us-west-2
For on-demand, provisioned throughput deployment, we want to prioritize backends in following orders based on the capacity.
- PT with small context length
- PT with large context length
- on-demand endpoint
Extend the envoy gateway configuration to support multiple backend priority levels which can be achieved by introducing the priority field in the backend configuration, allowing users to specify a numerical priority level.
[optional *Relevant Links*:]
https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/load_balancing/priority
https://github.com/envoyproxy/ai-gateway/issues/34
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.