agentscope-ai / agentscope-ai/agentscope-java

Harness close can return before SessionTree mirror finishes using distributed store

Open
#2,147 2 comments 0 reactions 0 assignees View on GitHub
area/harness bug
Dominant language
Java
Stars
5.6k
Forks
1.3k
Avg merge
4d 12h
Merged PRs (30d)
77

Description

## Environment

- AgentScope Java source commit: `e3a412ed2cc944e401da861c8d5e464b967724e9`
- HarnessAgent with `RedisDistributedStore` and `RemoteFilesystemSpec(IsolationScope.SESSION)`
- JedisPooled client owned by the deployment bootstrap

## Reproduction

1. Run a Harness turn that updates session history/offload on a Redis-backed remote filesystem.
2. Close the Harness runtime/agent.
3. Close the Jedis client after the agent close returns.
4. Observe the background `session-tree-mirror` thread.

## Actual

The test itself completes, then the background thread attempts `RemoteFilesystem.uploadFiles -> RedisStore.put` after the Jedis pool has already been closed:

```
Exception in thread "session-tree-mirror" redis.clients.jedis.exceptions.JedisException: Could not get a resource from the pool
...
Caused by: java.lang.IllegalStateException: Pool not open
...
at io.agentscope.harness.agent.memory.session.SessionTree.mirrorToFilesystem
at io.agentscope.harness.agent.memory.session.SessionTree.lambda$scheduleMirror$4
```

This means `HarnessAgent.close()` / workspace close does not provide a completion barrier for pending SessionTree mirror work. A deployment cannot know when it is safe to close the official distributed store client, and uncaught background failures can occur after tests/turns report success.

## Expected

The owner that schedules SessionTree mirror work should expose deterministic lifecycle semantics:

1. stop accepting new mirror jobs during close;
2. drain/await already accepted jobs, or cancel them with an observable terminal result;
3. close the mirror executor before `HarnessAgent.close()` returns;
4. surface mirror failures through a typed event/close failure or observability hook rather than an uncaught background exception;
5. make repeated close idempotent;
6. document the safe ordering for HarnessAgent, WorkspaceManager and DistributedStore/client close.

## Why upstream

Session history mirroring and its executor are Harness runtime facts. Product hosts should not sleep, poll private executors, keep Redis clients alive indefinitely, or add a second session mirror lifecycle manager.

A regression test should use a blocking BaseStore/RemoteFilesystem write, call agent close concurrently, and verify close does not return until the accepted mirror is drained/cancelled and no write occurs after store close.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.