agent-substrate / agent-substrate/substrate

Identify what metrics we should use for HPA-based workerpool autoscaling

Abierto
#671 4 comentarios 0 reacciones 1 asignado Reclamado por @shrutiyam-glitch Ver en GitHub
area/observability area/scheduling kind/design kind/feature
Lenguaje dominante
Go
Estrellas
1.8k
Forks
316
Merge medio
2 d 43 min
PR fusionados (30 d)
287

Descripción

Today, the HPA demo is configured to autoscale a workerpool based on `ate.workerpool.workers` metric: https://github.com/agent-substrate/substrate/blob/e70e1f21f85e20844548a521231868bacefd55d4/demos/autoscaled-workerpool/prometheus-adapter.yaml#L47-L57

`ate.workerpool.workers` counts workers by state, so it's clamped by capacity: eg, when a pool of 5 is fully assigned it reports 5, whether one actor is waiting or a thousand are.

Scaling on this metric alone may be fine for the steady state, but we'll be blind to incoming demand that greatly exceeds the pool capacity.

`atenet.router.parking.active` is a demand sign, but it doesn't know what specific workerpool the demand is for . A control plane equivalent metric doesn't exist.

One option is to autoscale on a new `ate.workerpool.demand` (filtered on `pool_exhausted` blocked reason) metric that counts actors that want a slot:

| attribute | values |
| --- | --- |
| `ate.workerpool.name` | pool name; empty when nothing in the fleet was eligible |
| `ate.sandbox.class` | `gvisor` / `microvm` |
| `ate.demand.state` | `resident` (holding a slot now) / `blocked` (rejected/still retrying) |
| `ate.blocked.reason` | `none`, `pool_exhausted`, `node_locality`, `no_eligible_pool`, `unknown` |

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.