agentscope-ai / agentscope-ai/agentscope-java

Harness close can return before SessionTree mirror finishes using distributed store

Abierto
#2,147 2 comentarios 0 reacciones 0 asignados Ver en GitHub
area/harness bug
Lenguaje dominante
Java
Estrellas
5.6k
Forks
1.3k
Merge medio
4 d 12 h
PR fusionados (30 d)
77

Descripción

## Environment

- AgentScope Java source commit: `e3a412ed2cc944e401da861c8d5e464b967724e9`
- HarnessAgent with `RedisDistributedStore` and `RemoteFilesystemSpec(IsolationScope.SESSION)`
- JedisPooled client owned by the deployment bootstrap

## Reproduction

1. Run a Harness turn that updates session history/offload on a Redis-backed remote filesystem.
2. Close the Harness runtime/agent.
3. Close the Jedis client after the agent close returns.
4. Observe the background `session-tree-mirror` thread.

## Actual

The test itself completes, then the background thread attempts `RemoteFilesystem.uploadFiles -> RedisStore.put` after the Jedis pool has already been closed:

```
Exception in thread "session-tree-mirror" redis.clients.jedis.exceptions.JedisException: Could not get a resource from the pool
...
Caused by: java.lang.IllegalStateException: Pool not open
...
at io.agentscope.harness.agent.memory.session.SessionTree.mirrorToFilesystem
at io.agentscope.harness.agent.memory.session.SessionTree.lambda$scheduleMirror$4
```

This means `HarnessAgent.close()` / workspace close does not provide a completion barrier for pending SessionTree mirror work. A deployment cannot know when it is safe to close the official distributed store client, and uncaught background failures can occur after tests/turns report success.

## Expected

The owner that schedules SessionTree mirror work should expose deterministic lifecycle semantics:

1. stop accepting new mirror jobs during close;
2. drain/await already accepted jobs, or cancel them with an observable terminal result;
3. close the mirror executor before `HarnessAgent.close()` returns;
4. surface mirror failures through a typed event/close failure or observability hook rather than an uncaught background exception;
5. make repeated close idempotent;
6. document the safe ordering for HarnessAgent, WorkspaceManager and DistributedStore/client close.

## Why upstream

Session history mirroring and its executor are Harness runtime facts. Product hosts should not sleep, poll private executors, keep Redis clients alive indefinitely, or add a second session mirror lifecycle manager.

A regression test should use a blocking BaseStore/RemoteFilesystem write, call agent close concurrently, and verify close does not return until the accepted mirror is drained/cancelled and no write occurs after store close.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.