microsoft / microsoft/agent-framework
Handling partial failure in a superstep
Open
@TaoChenOSU is already working on this.
Since Dec 5, 2025.
.NET
python
question
workflows
- Dominant language
- Python
- Stars
- 13.6k
- Forks
- 2.3k
- Avg merge
- 2d 45m
- Merged PRs (30d)
- 358
Description
Today, if an executor fails in a superstep while others succeed, we can only resume from the most recent superstep and rerun all executors.
It'd be great if we can create a checkpoint even if a superstep succeeds partially. That way, users can resume from that checkpoint and only rerun the executors that failed, saving costs and token consumption.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.