agent-substrate / agent-substrate/substrate

Prefer workers with cached actor images during scheduling

未关闭
#276 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
area/api-machinery area/scheduling kind/feature
主要语言
Go
星标
1.8k
派生
316
平均合并
2 天 43 分钟
30 天内合并 PR
287

描述

## Description

Atelet maintains a node-local image cache, but the scheduler does not know which images are cached on each node. As a result, an actor may be assigned to a cold node and pull its images again, even when another eligible node already has them cached.

Image preparation happens synchronously during `ResumeActor`, so unnecessary pulls increase actor startup latency and registry traffic.

## Proposal

- Let atelet report its cached image digests to the control plane.
- Store cache information by node.
- After applying existing constraints, prefer available workers whose node has all required actor images cached.
- Fall back to the current random selection when no cache hit is available.

Local snapshot availability remains a hard constraint. Image cache affinity is applied only among nodes that can restore the selected snapshot.

A general scoring framework is not required: use cache-hit preference with fallback.

## Expected Outcome

Reduce unnecessary image pulls and improve Actor Resume latency and consistency.

## Related Issues

- #166 and #228 optimize node-local image/rootfs reuse.
- #135 discusses locality-aware scheduling for cached snapshots.
- #233 tracks cold image pulls exceeding the `ResumeActor` deadline.

This proposal complements those efforts by making worker selection aware of node-local image cache availability.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。