agentscope-ai / agentscope-ai/agentscope-java

Harness close can return before SessionTree mirror finishes using distributed store

未關閉
#2,147 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
area/harness bug
主要語言
Java
星號
5.6k
分支
1.3k
平均合併
4 天 12 小時
30 天內合併 PR
77

描述

## Environment

- AgentScope Java source commit: `e3a412ed2cc944e401da861c8d5e464b967724e9`
- HarnessAgent with `RedisDistributedStore` and `RemoteFilesystemSpec(IsolationScope.SESSION)`
- JedisPooled client owned by the deployment bootstrap

## Reproduction

1. Run a Harness turn that updates session history/offload on a Redis-backed remote filesystem.
2. Close the Harness runtime/agent.
3. Close the Jedis client after the agent close returns.
4. Observe the background `session-tree-mirror` thread.

## Actual

The test itself completes, then the background thread attempts `RemoteFilesystem.uploadFiles -> RedisStore.put` after the Jedis pool has already been closed:

```
Exception in thread "session-tree-mirror" redis.clients.jedis.exceptions.JedisException: Could not get a resource from the pool
...
Caused by: java.lang.IllegalStateException: Pool not open
...
at io.agentscope.harness.agent.memory.session.SessionTree.mirrorToFilesystem
at io.agentscope.harness.agent.memory.session.SessionTree.lambda$scheduleMirror$4
```

This means `HarnessAgent.close()` / workspace close does not provide a completion barrier for pending SessionTree mirror work. A deployment cannot know when it is safe to close the official distributed store client, and uncaught background failures can occur after tests/turns report success.

## Expected

The owner that schedules SessionTree mirror work should expose deterministic lifecycle semantics:

1. stop accepting new mirror jobs during close;
2. drain/await already accepted jobs, or cancel them with an observable terminal result;
3. close the mirror executor before `HarnessAgent.close()` returns;
4. surface mirror failures through a typed event/close failure or observability hook rather than an uncaught background exception;
5. make repeated close idempotent;
6. document the safe ordering for HarnessAgent, WorkspaceManager and DistributedStore/client close.

## Why upstream

Session history mirroring and its executor are Harness runtime facts. Product hosts should not sleep, poll private executors, keep Redis clients alive indefinitely, or add a second session mirror lifecycle manager.

A regression test should use a blocking BaseStore/RemoteFilesystem write, call agent close concurrently, and verify close does not return until the accepted mirror is drained/cancelled and no write occurs after store close.

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。