agent-substrate / agent-substrate/substrate

Should ateom share network access with the actor sandbox?

未关闭
#240 4 条评论 0 个 reaction 已指派 1 人 已被 @bowei 认领 在 GitHub 查看
area/network area/node kind/feature
主要语言
Go
星标
1.8k
派生
316
平均合并
2 天 43 分钟
30 天内合并 PR
287

描述

Follow-up for discussion from #122

AFAICT: We're not actually using the otel trace export ... we export to localhost, so it's effectively a no-op currently. ateom doesn't actually use the network meaningfully.

While the traces would be nice to have, we could have atelet relay these (e.g. we could use another unix socket for ateom traces => atelet), which is in-line with how we're trending for scheduling (aggregate at the atelet, avoid every worker talking directly to ate-apiserver) but repeated for tracing and the trace collector.

So far there doesn't seem to be any other use-case for inbound connections to the ateom.

I'm not sure there's a strong case for ateom to intrude on the worker's network at all, I could see potentially streaming snapshots but I think we're moving away from always pushing these over the network anyhow (https://github.com/agent-substrate/substrate/issues/119, https://github.com/agent-substrate/substrate/pull/227) so we'll probably have some local cache in atelet that we need to pull through anyhow.

It would be nice to just forward all traffic on all protocols from the pod IP to the interior sandbox and avoid users having to configure explicit actor ports, and more importantly avoid conflicts with any ports actors might desire.

I feel like this would've made more sense if we didn't think we'd be writing snapshots to local disk commonly (to avoid too much traffic on the nic) and if we thought we were going to eliminate atelet (#128 , currently I disagree and I think atelet will be critical for aggregating both scheduling and images / snapshots / ...), but even then debugging it may be ... confusing.

cc @thockin @bowei @EItanya @ahmedtd

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。