agentscope-ai / agentscope-ai/agentscope-java

Harness close can return before SessionTree mirror finishes using distributed store

Đang mở
#2,147 2 bình luận 0 reaction 0 người được giao Xem trên GitHub
area/harness bug
Ngôn ngữ chính
Java
Star
5.6k
Fork
1.3k
Merge trung bình
4 ngày 12 giờ
Pull request đã merge (30 ngày)
77

Mô tả

## Environment

- AgentScope Java source commit: `e3a412ed2cc944e401da861c8d5e464b967724e9`
- HarnessAgent with `RedisDistributedStore` and `RemoteFilesystemSpec(IsolationScope.SESSION)`
- JedisPooled client owned by the deployment bootstrap

## Reproduction

1. Run a Harness turn that updates session history/offload on a Redis-backed remote filesystem.
2. Close the Harness runtime/agent.
3. Close the Jedis client after the agent close returns.
4. Observe the background `session-tree-mirror` thread.

## Actual

The test itself completes, then the background thread attempts `RemoteFilesystem.uploadFiles -> RedisStore.put` after the Jedis pool has already been closed:

```
Exception in thread "session-tree-mirror" redis.clients.jedis.exceptions.JedisException: Could not get a resource from the pool
...
Caused by: java.lang.IllegalStateException: Pool not open
...
at io.agentscope.harness.agent.memory.session.SessionTree.mirrorToFilesystem
at io.agentscope.harness.agent.memory.session.SessionTree.lambda$scheduleMirror$4
```

This means `HarnessAgent.close()` / workspace close does not provide a completion barrier for pending SessionTree mirror work. A deployment cannot know when it is safe to close the official distributed store client, and uncaught background failures can occur after tests/turns report success.

## Expected

The owner that schedules SessionTree mirror work should expose deterministic lifecycle semantics:

1. stop accepting new mirror jobs during close;
2. drain/await already accepted jobs, or cancel them with an observable terminal result;
3. close the mirror executor before `HarnessAgent.close()` returns;
4. surface mirror failures through a typed event/close failure or observability hook rather than an uncaught background exception;
5. make repeated close idempotent;
6. document the safe ordering for HarnessAgent, WorkspaceManager and DistributedStore/client close.

## Why upstream

Session history mirroring and its executor are Harness runtime facts. Product hosts should not sleep, poll private executors, keep Redis clients alive indefinitely, or add a second session mirror lifecycle manager.

A regression test should use a blocking BaseStore/RemoteFilesystem write, call agent close concurrently, and verify close does not return until the accepted mirror is drained/cancelled and no write occurs after store close.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.