envoyproxy / envoyproxy/gateway
Support for Backend priorities beyond active/passive levels
- Langage dominant
- Go
- Étoiles
- 3k
- Forks
- 864
- Merge moyen
- 2 j 2 h
- PR mergées (30 j)
- 140
Description
*Description*:
Currently, Envoy Gateway supports two levels of backend priorities: active and passive. While this is sufficient for basic use cases, it limits more advanced traffic management scenarios for GenAI traffic where multiple levels of backend priorities are required for multiple regions, model providers, on-demand and provisioned throughput with different context length.
For multi-region deployment, we want to prioritize backends in following orders to achieve higher SLO target.
- us-east-1
- us-east-2
- us-west-2
For on-demand, provisioned throughput deployment, we want to prioritize backends in following orders based on the capacity.
- PT with small context length
- PT with large context length
- on-demand endpoint
Extend the envoy gateway configuration to support multiple backend priority levels which can be achieved by introducing the priority field in the backend configuration, allowing users to specify a numerical priority level.
[optional *Relevant Links*:]
https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/upstream/load_balancing/priority
https://github.com/envoyproxy/ai-gateway/issues/34
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Évaluation
Cette issue n'a pas encore été évaluée.