aws / aws/aws-durable-execution-sdk-python
[Feature]: Reject new durable operations under orphaned map/parallel branches
- Dominant language
- Python
- Stars
- 53
- Forks
- 25
- Avg merge
- 1d 17h
- Merged PRs (30d)
- 37
Description
## What would you like?
Reject newly created durable operations whose parent context has already been
marked orphaned after an early-completing `map` or `parallel` operation.
Today, `_mark_orphans()` marks the descendants known when the parent context
completes. If an in-flight orphaned branch subsequently creates a new durable
operation, `ExecutionState.create_checkpoint()` adds the new operation to
`_parent_to_children` and checks only whether the new operation's own ID is in
`_parent_done`.
Because the new operation did not exist when `_mark_orphans()` took its
snapshot, its ID is not in `_parent_done`, so its checkpoint can be accepted
even though its `parent_id` is already orphaned.
The orphaned branch's own terminal result is normally rejected later, and the
parent `BatchResult` remains consistent across replay:
- Normal payloads replay the parent's serialized result.
- `ReplayChildren` summaries preserve the branch as `STARTED` through
`startedIndexes`.
This is therefore an improvement to lifecycle enforcement rather than a
customer-visible replay-result correctness bug. It would still prevent
unnecessary durable work, side effects, and history entries from being created
inside a branch whose parent has already completed.
This issue is separate from #640, which covers the race where an already-known
branch passes orphan validation before parent completion and enqueues its own
terminal result afterward.
## Possible Implementation
When admitting an operation checkpoint under `_parent_done_lock`, reject the
operation when either:
```python
operation_update.operation_id in self._parent_done
```
or:
```python
operation_update.parent_id in self._parent_done
```
If a parent can be indirectly orphaned without appearing directly in
`_parent_done`, propagate orphan state when registering the new child or walk
the known parent chain under the same lock.
The check should also be repeated in the atomic validation-and-enqueue section
proposed by #640 so parent completion cannot race queue insertion.
Suggested tests:
- A new step created under an already orphaned nested map iteration.
- A new invoke/callback/wait created under an orphaned parallel branch.
- The equivalent behavior for `NestingType.FLAT`.
- The new operation never reaches the checkpoint service.
- Parent results remain identical on first execution and replay for normal and
`ReplayChildren` payloads.
## Is this a breaking change?
No. The operation is already outside the lifetime of its completed parent
context and its enclosing branch cannot contribute a terminal result to the
completed batch.
## Does this require an RFC?
No.
## Additional Context
A deterministic state-level reproduction shows:
```text
branch orphaned: True
late descendant rejected: NO (accepted)
late descendant orphaned: False
```
Relevant code:
- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `ExecutionState.create_checkpoint()`
- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `_mark_orphans()`
Observed against SDK version `1.8.0`, repository HEAD
`7ac7acc6a7dae231f2abbb8e37f9780cc9b89af0`, using Python 3.14.
Contributor guide
Assessment
This issue has not been assessed yet.