agentscope-ai / agentscope-ai/agentscope-java
Distributed replicas can lose concurrent turns for the same agent session
- Lingua principale
- Java
- Stelle
- 5.6k
- Fork
- 1.3k
- Merge medio
- 4g 12h
- PR unite (30g)
- 77
Descrizione
## Environment
- AgentScope Java source commit: `e3a412ed2cc944e401da861c8d5e464b967724e9`
- Two independent `HarnessAgent` / `ReActAgent` instances
- Two independent Jedis clients sharing one official `RedisDistributedStore`
- Same `(userId, sessionId)`
- Different JVM/runtime replica identity
## Reproduction
1. Build two Harness runtimes backed by the same Redis distributed store and session.
2. Use a model barrier so both calls enter model execution after both agents have loaded the same initial AgentState.
3. Submit a distinct user turn through each runtime concurrently.
4. Release both model calls and wait for both agent calls to complete successfully.
5. Reload `agent_state` from Redis.
## Actual
Both calls return successfully, but persisted AgentState contains only one of the two turns. The other turn is lost.
The current same-session guards are process-local:
- `AgentBase.callGates` serializes by slot only inside one agent instance/JVM.
- Harness `SessionTurnGate` uses an in-memory `ConcurrentHashMap`.
- `AgentStateStore.save(...)` is unconditional full replacement.
- `RedisAgentStateStore.save(...)` performs a plain set without version/CAS/fencing.
Reloading AgentState at the start of each call helps replica drift between sequential calls, but cannot prevent two replicas from concurrently loading the same version and overwriting each other.
## Expected
AgentScope should expose an official distributed same-session serialization contract, for example:
- a lease/claim with TTL, owner token and fencing token around the complete agent turn; or
- versioned AgentState load/save with bounded conflict handling that is safe for side-effecting tools; or
- an official distributed session router/guard abstraction backed by Redis/JDBC.
Required behavior:
1. same `(userId, sessionId)` across replicas cannot execute conflicting turns concurrently;
2. lease expiry/crash recovery does not allow a stale worker to commit after takeover;
3. cancellation and HITL suspension release or transfer ownership deterministically;
4. different sessions remain parallel;
5. state, tool execution and emitted terminal events agree on one winning fenced turn;
6. fresh-node continuation is supported without product-host locks.
## Why upstream
Agent execution serialization, AgentState persistence and HITL/tool side effects are Core/Harness runtime facts. Product hosts should not build a second Redis lock, queue or session coordinator around AgentScope.
A regression test can use two independent Redis clients and two Harness runtimes with a model barrier, then assert both turns are retained or one is explicitly rejected before model/tool execution.
Guida per i contributori
Apri la guida per i contributori
Valutazione
Questa issue non è ancora stata valutata.