aws / aws/aws-durable-execution-sdk-python

[Feature]: Reject new durable operations under orphaned map/parallel branches

未关闭
#641 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
enhancement
主要语言
Python
星标
53
派生
25
平均合并
1 天 19 小时
30 天内合并 PR
40

描述

## What would you like?

Reject newly created durable operations whose parent context has already been
marked orphaned after an early-completing `map` or `parallel` operation.

Today, `_mark_orphans()` marks the descendants known when the parent context
completes. If an in-flight orphaned branch subsequently creates a new durable
operation, `ExecutionState.create_checkpoint()` adds the new operation to
`_parent_to_children` and checks only whether the new operation's own ID is in
`_parent_done`.

Because the new operation did not exist when `_mark_orphans()` took its
snapshot, its ID is not in `_parent_done`, so its checkpoint can be accepted
even though its `parent_id` is already orphaned.

The orphaned branch's own terminal result is normally rejected later, and the
parent `BatchResult` remains consistent across replay:

- Normal payloads replay the parent's serialized result.
- `ReplayChildren` summaries preserve the branch as `STARTED` through
`startedIndexes`.

This is therefore an improvement to lifecycle enforcement rather than a
customer-visible replay-result correctness bug. It would still prevent
unnecessary durable work, side effects, and history entries from being created
inside a branch whose parent has already completed.

This issue is separate from #640, which covers the race where an already-known
branch passes orphan validation before parent completion and enqueues its own
terminal result afterward.

## Possible Implementation

When admitting an operation checkpoint under `_parent_done_lock`, reject the
operation when either:

```python
operation_update.operation_id in self._parent_done
```

or:

```python
operation_update.parent_id in self._parent_done
```

If a parent can be indirectly orphaned without appearing directly in
`_parent_done`, propagate orphan state when registering the new child or walk
the known parent chain under the same lock.

The check should also be repeated in the atomic validation-and-enqueue section
proposed by #640 so parent completion cannot race queue insertion.

Suggested tests:

- A new step created under an already orphaned nested map iteration.
- A new invoke/callback/wait created under an orphaned parallel branch.
- The equivalent behavior for `NestingType.FLAT`.
- The new operation never reaches the checkpoint service.
- Parent results remain identical on first execution and replay for normal and
`ReplayChildren` payloads.

## Is this a breaking change?

No. The operation is already outside the lifetime of its completed parent
context and its enclosing branch cannot contribute a terminal result to the
completed batch.

## Does this require an RFC?

No.

## Additional Context

A deterministic state-level reproduction shows:

```text
branch orphaned: True
late descendant rejected: NO (accepted)
late descendant orphaned: False
```

Relevant code:

- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `ExecutionState.create_checkpoint()`
- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `_mark_orphans()`

Observed against SDK version `1.8.0`, repository HEAD
`7ac7acc6a7dae231f2abbb8e37f9780cc9b89af0`, using Python 3.14.

贡献指南

打开贡献指南

调研方向

阅读 packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py,从 ExecutionState.create_checkpoint() 和 _mark_orphans() 开始。重现延迟后代的情况,然后为 nested map、parallel、invoke/callback/wait 和 FLAT 分支添加测试。当孤立操作在创建检查点之前被拒绝,同时首次执行和 replay 中的 parent 结果保持一致时,即表示完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
backend, distributed-systems
Issue 类型
功能
难度
4/5
预计耗时
3-5 天
活跃度
冷清
描述清晰度
描述清楚
新手友好度
52/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。