agent-substrate / agent-substrate/substrate

Identify what metrics we should use for HPA-based workerpool autoscaling

Offen
#671 4 Kommentare 0 Reaktionen 1 zugewiesene Person Beansprucht von @shrutiyam-glitch Auf GitHub ansehen
area/observability area/scheduling kind/design kind/feature
Vorherrschende Sprache
Go
Sterne
1.8k
Forks
316
Ø Merge
2 T. 43 Min.
Gemergte PRs (30 T.)
287

Beschreibung

Today, the HPA demo is configured to autoscale a workerpool based on `ate.workerpool.workers` metric: https://github.com/agent-substrate/substrate/blob/e70e1f21f85e20844548a521231868bacefd55d4/demos/autoscaled-workerpool/prometheus-adapter.yaml#L47-L57

`ate.workerpool.workers` counts workers by state, so it's clamped by capacity: eg, when a pool of 5 is fully assigned it reports 5, whether one actor is waiting or a thousand are.

Scaling on this metric alone may be fine for the steady state, but we'll be blind to incoming demand that greatly exceeds the pool capacity.

`atenet.router.parking.active` is a demand sign, but it doesn't know what specific workerpool the demand is for . A control plane equivalent metric doesn't exist.

One option is to autoscale on a new `ate.workerpool.demand` (filtered on `pool_exhausted` blocked reason) metric that counts actors that want a slot:

| attribute | values |
| --- | --- |
| `ate.workerpool.name` | pool name; empty when nothing in the fleet was eligible |
| `ate.sandbox.class` | `gvisor` / `microvm` |
| `ate.demand.state` | `resident` (holding a slot now) / `blocked` (rejected/still retrying) |
| `ate.blocked.reason` | `none`, `pool_exhausted`, `node_locality`, `no_eligible_pool`, `unknown` |

Beitragsleitfaden

Beitragsleitfaden öffnen

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.