aws / aws/aws-durable-execution-sdk-python

[Feature]: Reject new durable operations under orphaned map/parallel branches

オープン
#641 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
enhancement
主要言語
Python
スター
53
フォーク
25
平均マージ
1日 17時間
マージ済み PR(30日)
37

説明

## What would you like?

Reject newly created durable operations whose parent context has already been
marked orphaned after an early-completing `map` or `parallel` operation.

Today, `_mark_orphans()` marks the descendants known when the parent context
completes. If an in-flight orphaned branch subsequently creates a new durable
operation, `ExecutionState.create_checkpoint()` adds the new operation to
`_parent_to_children` and checks only whether the new operation's own ID is in
`_parent_done`.

Because the new operation did not exist when `_mark_orphans()` took its
snapshot, its ID is not in `_parent_done`, so its checkpoint can be accepted
even though its `parent_id` is already orphaned.

The orphaned branch's own terminal result is normally rejected later, and the
parent `BatchResult` remains consistent across replay:

- Normal payloads replay the parent's serialized result.
- `ReplayChildren` summaries preserve the branch as `STARTED` through
`startedIndexes`.

This is therefore an improvement to lifecycle enforcement rather than a
customer-visible replay-result correctness bug. It would still prevent
unnecessary durable work, side effects, and history entries from being created
inside a branch whose parent has already completed.

This issue is separate from #640, which covers the race where an already-known
branch passes orphan validation before parent completion and enqueues its own
terminal result afterward.

## Possible Implementation

When admitting an operation checkpoint under `_parent_done_lock`, reject the
operation when either:

```python
operation_update.operation_id in self._parent_done
```

or:

```python
operation_update.parent_id in self._parent_done
```

If a parent can be indirectly orphaned without appearing directly in
`_parent_done`, propagate orphan state when registering the new child or walk
the known parent chain under the same lock.

The check should also be repeated in the atomic validation-and-enqueue section
proposed by #640 so parent completion cannot race queue insertion.

Suggested tests:

- A new step created under an already orphaned nested map iteration.
- A new invoke/callback/wait created under an orphaned parallel branch.
- The equivalent behavior for `NestingType.FLAT`.
- The new operation never reaches the checkpoint service.
- Parent results remain identical on first execution and replay for normal and
`ReplayChildren` payloads.

## Is this a breaking change?

No. The operation is already outside the lifetime of its completed parent
context and its enclosing branch cannot contribute a terminal result to the
completed batch.

## Does this require an RFC?

No.

## Additional Context

A deterministic state-level reproduction shows:

```text
branch orphaned: True
late descendant rejected: NO (accepted)
late descendant orphaned: False
```

Relevant code:

- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `ExecutionState.create_checkpoint()`
- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `_mark_orphans()`

Observed against SDK version `1.8.0`, repository HEAD
`7ac7acc6a7dae231f2abbb8e37f9780cc9b89af0`, using Python 3.14.

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.pyを読み、ExecutionState.create_checkpoint()と_mark_orphans()から始めます。遅れて生成される子孫のケースを再現し、その後、nested map、parallel、invoke/callback/wait、FLATブランチのテストを追加します。初回実行とreplayでParentの結果が同一のまま、チェックポイント作成前に孤立した操作が拒否されれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
backend, distributed-systems
issue の種類
機能追加
難易度
4/5
見積もり時間
3〜5日
活発さ
静か
明瞭さ
明確に書かれている
初心者へのやさしさ
52/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。