agentscope-ai / agentscope-ai/agentscope-java

[Feature]: Expose typed model retry and fallback attempt events

Đang mở
#2,136 1 bình luận 0 reaction 0 người được giao Xem trên GitHub
area/core/model enhancement
Ngôn ngữ chính
Java
Star
5.6k
Fork
1.3k
Merge trung bình
4 ngày 12 giờ
Pull request đã merge (30 ngày)
77

Mô tả

## Motivation

AgentScope 2.0 supports model retry and fallback through `ExecutionConfig` / `ModelUtils` and `ReActAgent.Builder.maxRetries(...)` / `fallbackModel(...)`. The runtime behavior works, but `streamEvents()` currently exposes only one aggregate `MODEL_CALL_START` and `MODEL_CALL_END` boundary for the entire operation.

A product host therefore cannot answer, from typed runtime facts:

- which provider/model served the successful response;
- whether the primary was retried;
- whether fallback was activated;
- how many attempts occurred;
- whether the failure was rate-limit, timeout, authentication, network, or provider 5xx;
- which attempt consumed the reported usage/latency.

Parsing logs or synthesizing attempt facts in the host would create a second model-runtime truth source.

## Current behavior

`ReActAgent#modelCallStream(...)` emits:

```text
MODEL_CALL_START -> content block events -> MODEL_CALL_END
```

`modelForCall()` wraps primary/fallback internally. If the primary errors before its first signal, the wrapper switches to the fallback, but no typed event identifies that transition. Provider-level retry is applied inside model transport through `ModelUtils.applyTimeoutAndRetry(...)`, which is also not visible through `AgentEvent`.

A direct contract with a failing primary and successful fallback observes exactly one `MODEL_CALL_START`, one `MODEL_CALL_END`, and no retry/fallback event, even though the fallback completed the turn.

## Proposed contract

Please expose a typed attempt lifecycle owned by Core. Exact naming is open for discussion; one possible shape is:

```text
MODEL_ATTEMPT_START
MODEL_ATTEMPT_FAILED
MODEL_FALLBACK_ACTIVATED
MODEL_ATTEMPT_END
```

Suggested fields:

- logical model-call / reply id;
- attempt index and configured max attempts;
- provider id and model id/name;
- primary vs fallback role;
- typed failure category (`rate_limit`, `timeout`, `authentication`, `authorization`, `network`, `provider_5xx`, `invalid_request`, `unknown`);
- retryable flag and next action (`retry`, `fallback`, `fail`);
- sanitized status/error code without credential or response-body leakage;
- per-attempt latency and usage when available.

The existing aggregate `MODEL_CALL_START/END` can remain backward compatible around the full logical call.

## Required properties

- Events must come from Core/model transport, not product-host log parsing.
- No API keys, authorization headers, raw credential values, or sensitive response bodies.
- Works for both direct provider retries and `fallbackModel` transitions.
- Failed attempts do not create assistant messages or replay completed tool calls.
- Same behavior through `call()` and `streamEvents()` lifecycle.
- Provider implementations can map native errors into the shared failure taxonomy.

## Validation proposal

Tests should cover:

1. 429 and 5xx retry, followed by success.
2. 401/403 non-retryable classification.
3. timeout/network retry.
4. primary exhaustion followed by fallback success.
5. fallback failure.
6. fallback after a completed tool result without replaying the tool.
7. HITL approval continuation followed by fallback.
8. secret redaction in serialized events.

I have downstream contracts confirming fallback continuity, single execution of side-effecting tools, and the current aggregate-only event gap. I can prepare a focused PR after maintainers confirm the preferred event shape and whether retry events should originate in `ModelUtils`, middleware, or `ReActAgent`.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.