agentscope-ai / agentscope-ai/agentscope-java

[Bug]: WaitAsyncResultsTool waits on the inbox notification chain instead of task state, stalling the main agent when any link breaks

Abierto
#2,791 0 comentarios 0 reacciones 0 asignados Ver en GitHub
bug
Lenguaje dominante
Java
Estrellas
5.6k
Forks
1.3k
Merge medio
4 d 12 h
PR fusionados (30 d)
77

Descripción

**Describe the bug**

`WaitAsyncResultsTool` builds its waiting semantics on an indirect, multi-hop chain rather than
observing task state directly:

```
subtask completes -> TaskRepository completionCallback -> MessageBus inbox push
-> InboxMiddleware injects before next LLM call -> LLM sees the result
```

The tool itself only polls the inbox every 3s. If any link in that chain fails to deliver, the
tool never observes the fact that the task has actually finished — and it never cross-checks the
real task status in `TaskRepository`.

Combined with the `MAX_CONSECUTIVE_EMPTY_WAITS` rejection, the consequence escalates from "waits
too long" to "hard stall": after two consecutive empty waits the tool refuses to wait any longer,
while the main agent has neither received results nor been given any other way to make progress.

**To Reproduce**

1. You code

```java
HarnessAgent supervisor = HarnessAgent.builder()
.model(model)
.messageBus(messageBus)
.subagentFactory("worker", name -> workerAgent)
.build();

// Spawn one or more async subagents (background task mode), then let the supervisor
// call wait_async_results to collect their results.
supervisor.streamEvents(new UserMessage(promptThatSpawnsThenWaits), ctx).subscribe();
```

2. How to execute

Run the flow, and make any one link of the injection chain fail. Any of these is enough:

- another `WorkspaceTaskRepository.setCompletionCallback` call replaces the callback, so the
inbox push no longer happens (later registration overwrites the earlier one);
- the task reaches a terminal state between `putTask` and callback registration (race), so the
completion event is lost;
- the inbox message is consumed but injection lands after the round that needed it (injector race).

Then compare the subagent's real state with what the tool reports.

3. See error

`TaskRepository` shows the task as `COMPLETED`, the inbox is empty, and the tool keeps waiting
until timeout, then refuses to wait again. The LLM never receives the results.

**Expected behavior**

Once a task has actually reached a terminal state, `wait_async_results` should observe it and hand
the result to the LLM. Even if the inbox notification side-channel fails, the tool must not treat
"notification missing" as "task not finished".

**Error messages**

No exception is raised. The failure is silent, so the observable symptoms are:

- the tool blocks for the full 60s timeout although the task is already terminal;
- after two consecutive empty waits, waiting is refused via `MAX_CONSECUTIVE_EMPTY_WAITS` and the
conversation cannot proceed;
- when the inbox is empty but all tasks are terminal, the tool returns a promise such as
`results will be injected automatically`, which is never fulfilled in this scenario and actively
misleads the LLM into waiting further.

**Environment (please complete the following information):**

- AgentScope-Java Version: 2.0.0 (agentscope-harness 2.0.0)
- Java Version: 17
- OS: macos

**Additional context**

*Root cause.* The tool's wait condition is not the same thing as the semantics it means to express:

| | Intended semantics | What is actually observed |
|---|---|---|
| Goal | whether tasks reached a terminal state | task status in `TaskRepository` |
| Actual | whether the inbox has a message | a product of the notification chain |

Task status is the **authoritative fact** and is guaranteed to happen (unless the task itself
hangs). The inbox message is only a **derived notification** with several failure points along the
way. Observing the derived signal without checking the authoritative one makes "lost notification"
indistinguishable from "task not finished".

`MAX_CONSECUTIVE_EMPTY_WAITS` amplifies the defect. It assumes repeated empty waits mean the LLM
is abusing the wait, but when the notification chain is broken, repeated empty waits are a symptom
of the tool's own faulty observation — refusing to wait at that moment only pushes the LLM into a
dead end.

*Suggested fix.* Change the wait condition to poll `TaskRepository` for terminal task state
directly. `TaskRepository.listTasks(ctx, sessionId, null)` already returns all tasks with their
`TaskStatus`, and `TaskStatus.isTerminal()` already exists, so no new API is required. Once tasks
are terminal, return the result text as the tool's return value so result collection closes within
a single `call()` and no longer depends on cross-round injection. Inbox injection can stay as a
fast path, but should not be the only correctness dependency.

Secondary points:

1. Remove or reposition `MAX_CONSECUTIVE_EMPTY_WAITS`. If the wait observes real task state, a
terminal state is guaranteed to arrive and the rejection is meaningless; if the intent is to
guard against hung tasks, the timeout itself is the better mechanism.
2. When the inbox is empty and all tasks are terminal, do not return
`will be injected automatically`. Return the actual results, or state the current status
explicitly, instead of making a promise that will not be kept.

Affected classes:

- `io.agentscope.harness.agent.tool.WaitAsyncResultsTool`
- `io.agentscope.harness.agent.subagent.task.WorkspaceTaskRepository` (completion callback registration)

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.