aws / aws/aws-durable-execution-sdk-python
[Feature]: Reject new durable operations under orphaned map/parallel branches
- 主要语言
- Python
- 星标
- 53
- 派生
- 25
- 平均合并
- 1 天 19 小时
- 30 天内合并 PR
- 40
描述
## What would you like?
Reject newly created durable operations whose parent context has already been
marked orphaned after an early-completing `map` or `parallel` operation.
Today, `_mark_orphans()` marks the descendants known when the parent context
completes. If an in-flight orphaned branch subsequently creates a new durable
operation, `ExecutionState.create_checkpoint()` adds the new operation to
`_parent_to_children` and checks only whether the new operation's own ID is in
`_parent_done`.
Because the new operation did not exist when `_mark_orphans()` took its
snapshot, its ID is not in `_parent_done`, so its checkpoint can be accepted
even though its `parent_id` is already orphaned.
The orphaned branch's own terminal result is normally rejected later, and the
parent `BatchResult` remains consistent across replay:
- Normal payloads replay the parent's serialized result.
- `ReplayChildren` summaries preserve the branch as `STARTED` through
`startedIndexes`.
This is therefore an improvement to lifecycle enforcement rather than a
customer-visible replay-result correctness bug. It would still prevent
unnecessary durable work, side effects, and history entries from being created
inside a branch whose parent has already completed.
This issue is separate from #640, which covers the race where an already-known
branch passes orphan validation before parent completion and enqueues its own
terminal result afterward.
## Possible Implementation
When admitting an operation checkpoint under `_parent_done_lock`, reject the
operation when either:
```python
operation_update.operation_id in self._parent_done
```
or:
```python
operation_update.parent_id in self._parent_done
```
If a parent can be indirectly orphaned without appearing directly in
`_parent_done`, propagate orphan state when registering the new child or walk
the known parent chain under the same lock.
The check should also be repeated in the atomic validation-and-enqueue section
proposed by #640 so parent completion cannot race queue insertion.
Suggested tests:
- A new step created under an already orphaned nested map iteration.
- A new invoke/callback/wait created under an orphaned parallel branch.
- The equivalent behavior for `NestingType.FLAT`.
- The new operation never reaches the checkpoint service.
- Parent results remain identical on first execution and replay for normal and
`ReplayChildren` payloads.
## Is this a breaking change?
No. The operation is already outside the lifetime of its completed parent
context and its enclosing branch cannot contribute a terminal result to the
completed batch.
## Does this require an RFC?
No.
## Additional Context
A deterministic state-level reproduction shows:
```text
branch orphaned: True
late descendant rejected: NO (accepted)
late descendant orphaned: False
```
Relevant code:
- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `ExecutionState.create_checkpoint()`
- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `_mark_orphans()`
Observed against SDK version `1.8.0`, repository HEAD
`7ac7acc6a7dae231f2abbb8e37f9780cc9b89af0`, using Python 3.14.
贡献指南
调研方向
阅读 packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py,从 ExecutionState.create_checkpoint() 和 _mark_orphans() 开始。重现延迟后代的情况,然后为 nested map、parallel、invoke/callback/wait 和 FLAT 分支添加测试。当孤立操作在创建检查点之前被拒绝,同时首次执行和 replay 中的 parent 结果保持一致时,即表示完成。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- backend, distributed-systems
- Issue 类型
- 功能
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 冷清
- 描述清晰度
- 描述清楚
- 新手友好度
- 52/100