aws / aws/aws-durable-execution-sdk-java
[Feature]: Exit gracefully with PENDING when a checkpoint response has no CheckpointToken
- 主要语言
- Java
- 星标
- 28
- 派生
- 11
- 平均合并
- 1 天 10 小时
- 30 天内合并 PR
- 45
描述
### What would you like?
When a checkpoint response arrives without a `CheckpointToken`, the SDK should treat it as a signal that the service will accept no further checkpoints from this invocation. The SDK should stop issuing checkpoints and end the invocation cleanly with `Status: PENDING`.
Why PENDING and not FAILED or a thrown error:
- A missing token does not mean the execution is finished. It means this invocation cannot make further progress. The invocation result must not claim the execution finished.
- A thrown error is an invocation failure. Lambda retries it, but the retry has no valid token and can do nothing useful. It also puts an error in customer logs for a condition the SDK understood.
- PENDING is already what the SDK returns for every suspend (wait, scheduled retry, pending callback). This is a suspend: the invocation is done for now and the execution continues later.
### Current behavior
`CheckpointManager.java` assigns the response token unconditionally:
```java
checkpointToken = response.checkpointToken();
```
So after a response with no token, `checkpointToken` is null. The next checkpoint call passes null to the client. The outcome is either a client-side parameter validation exception or a service `InvalidParameterValueException: Invalid checkpoint token` rejection; I have not verified which. A client-side exception is not an `AwsServiceException`, so `DurableApiErrorClassifier.classifyException` never sees it and it propagates as-is. A service rejection is classified per #705: with that fix it is a retryable invocation error and the handler throws; without it, the execution fails. In no case does the SDK return PENDING.
### Possible Implementation
1. After the checkpoint call, if `response.checkpointToken()` is null, mark the manager as suspended-by-service, stop issuing further checkpoints, and complete any pending pollers so no thread blocks on a checkpoint that will never happen.
2. Route the outcome through the existing suspension path that returns `InvocationStatus.PENDING` (the same path used by waits and callbacks).
3. In-flight operations are abandoned, not checkpointed. On the next invocation they replay. This is correct for AT_LEAST_ONCE steps. AT_MOST_ONCE steps whose START landed in the last accepted checkpoint will not run again, which is the defined semantics of AT_MOST_ONCE.
4. Unit test: a checkpoint response with a null token produces `PENDING`, no further checkpoint calls, and no thrown exception.
5. Testing module: the local runner needs a way to omit the token on a chosen checkpoint so customers can test their handlers against this path.
### Is this a breaking change?
No. The SDK's response set is unchanged; a new internal termination reason is added.
### Additional Context
- Prerequisite: #705 (message-case bug in stale-token classification). That fix is the fallback path whenever the missing token is not detected, so it should land first.
- Sibling issues in the JS and Python SDKs are linked in a comment below.
贡献指南
调研方向
从 CheckpointManager.java 开始,跟踪检查点响应处理以及返回 InvocationStatus.PENDING 的现有挂起路径。检查 DurableApiErrorClassifier 和测试模块的本地 runner,然后添加覆盖范围,证明 null token 会返回 PENDING、不再进行任何检查点调用且不会抛出异常。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- java
- 领域
- backend, testing
- Issue 类型
- 功能
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 活跃
- 描述清晰度
- 描述清楚
- 新手友好度
- 68/100