aws / aws/aws-durable-execution-sdk-java
[Feature]: Exit gracefully with PENDING when a checkpoint response has no CheckpointToken
- Dominant language
- Java
- Stars
- 28
- Forks
- 11
- Avg merge
- 1d 8h
- Merged PRs (30d)
- 47
Description
### What would you like?
When a checkpoint response arrives without a `CheckpointToken`, the SDK should treat it as a signal that the service will accept no further checkpoints from this invocation. The SDK should stop issuing checkpoints and end the invocation cleanly with `Status: PENDING`.
Why PENDING and not FAILED or a thrown error:
- A missing token does not mean the execution is finished. It means this invocation cannot make further progress. The invocation result must not claim the execution finished.
- A thrown error is an invocation failure. Lambda retries it, but the retry has no valid token and can do nothing useful. It also puts an error in customer logs for a condition the SDK understood.
- PENDING is already what the SDK returns for every suspend (wait, scheduled retry, pending callback). This is a suspend: the invocation is done for now and the execution continues later.
### Current behavior
`CheckpointManager.java` assigns the response token unconditionally:
```java
checkpointToken = response.checkpointToken();
```
So after a response with no token, `checkpointToken` is null. The next checkpoint call passes null to the client. The outcome is either a client-side parameter validation exception or a service `InvalidParameterValueException: Invalid checkpoint token` rejection; I have not verified which. A client-side exception is not an `AwsServiceException`, so `DurableApiErrorClassifier.classifyException` never sees it and it propagates as-is. A service rejection is classified per #705: with that fix it is a retryable invocation error and the handler throws; without it, the execution fails. In no case does the SDK return PENDING.
### Possible Implementation
1. After the checkpoint call, if `response.checkpointToken()` is null, mark the manager as suspended-by-service, stop issuing further checkpoints, and complete any pending pollers so no thread blocks on a checkpoint that will never happen.
2. Route the outcome through the existing suspension path that returns `InvocationStatus.PENDING` (the same path used by waits and callbacks).
3. In-flight operations are abandoned, not checkpointed. On the next invocation they replay. This is correct for AT_LEAST_ONCE steps. AT_MOST_ONCE steps whose START landed in the last accepted checkpoint will not run again, which is the defined semantics of AT_MOST_ONCE.
4. Unit test: a checkpoint response with a null token produces `PENDING`, no further checkpoint calls, and no thrown exception.
5. Testing module: the local runner needs a way to omit the token on a chosen checkpoint so customers can test their handlers against this path.
### Is this a breaking change?
No. The SDK's response set is unchanged; a new internal termination reason is added.
### Additional Context
- Prerequisite: #705 (message-case bug in stale-token classification). That fix is the fallback path whenever the missing token is not detected, so it should land first.
- Sibling issues in the JS and Python SDKs are linked in a comment below.
Contributor guide
Research direction
Start in CheckpointManager.java by tracing checkpoint response handling and the existing suspension path that returns InvocationStatus.PENDING. Review DurableApiErrorClassifier and the testing module's local runner, then add coverage showing that a null token returns PENDING, makes no further checkpoint calls, and throws no exception.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- backend, testing
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 68/100