aws / aws/aws-durable-execution-sdk-python

[Feature]: Reject new durable operations under orphaned map/parallel branches

Open
#641 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
53
Forks
25
Avg merge
1d 17h
Merged PRs (30d)
37

Description

## What would you like?

Reject newly created durable operations whose parent context has already been
marked orphaned after an early-completing `map` or `parallel` operation.

Today, `_mark_orphans()` marks the descendants known when the parent context
completes. If an in-flight orphaned branch subsequently creates a new durable
operation, `ExecutionState.create_checkpoint()` adds the new operation to
`_parent_to_children` and checks only whether the new operation's own ID is in
`_parent_done`.

Because the new operation did not exist when `_mark_orphans()` took its
snapshot, its ID is not in `_parent_done`, so its checkpoint can be accepted
even though its `parent_id` is already orphaned.

The orphaned branch's own terminal result is normally rejected later, and the
parent `BatchResult` remains consistent across replay:

- Normal payloads replay the parent's serialized result.
- `ReplayChildren` summaries preserve the branch as `STARTED` through
`startedIndexes`.

This is therefore an improvement to lifecycle enforcement rather than a
customer-visible replay-result correctness bug. It would still prevent
unnecessary durable work, side effects, and history entries from being created
inside a branch whose parent has already completed.

This issue is separate from #640, which covers the race where an already-known
branch passes orphan validation before parent completion and enqueues its own
terminal result afterward.

## Possible Implementation

When admitting an operation checkpoint under `_parent_done_lock`, reject the
operation when either:

```python
operation_update.operation_id in self._parent_done
```

or:

```python
operation_update.parent_id in self._parent_done
```

If a parent can be indirectly orphaned without appearing directly in
`_parent_done`, propagate orphan state when registering the new child or walk
the known parent chain under the same lock.

The check should also be repeated in the atomic validation-and-enqueue section
proposed by #640 so parent completion cannot race queue insertion.

Suggested tests:

- A new step created under an already orphaned nested map iteration.
- A new invoke/callback/wait created under an orphaned parallel branch.
- The equivalent behavior for `NestingType.FLAT`.
- The new operation never reaches the checkpoint service.
- Parent results remain identical on first execution and replay for normal and
`ReplayChildren` payloads.

## Is this a breaking change?

No. The operation is already outside the lifetime of its completed parent
context and its enclosing branch cannot contribute a terminal result to the
completed batch.

## Does this require an RFC?

No.

## Additional Context

A deterministic state-level reproduction shows:

```text
branch orphaned: True
late descendant rejected: NO (accepted)
late descendant orphaned: False
```

Relevant code:

- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `ExecutionState.create_checkpoint()`
- `packages/aws-durable-execution-sdk-python/src/aws_durable_execution_sdk_python/state.py`
in `_mark_orphans()`

Observed against SDK version `1.8.0`, repository HEAD
`7ac7acc6a7dae231f2abbb8e37f9780cc9b89af0`, using Python 3.14.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.