openworkflowdev / openworkflowdev/openworkflow
Flaky parallel crash-recovery test needs explicit step synchronization
@yqin0512 ya está trabajando en esto.
Desde el 13/9/2026.
- Lenguaje dominante
- TypeScript
- Estrellas
- 1.3k
- Forks
- 66
- Merge medio
- 1 d 1 h
- PR fusionados (30 d)
- 69
Descripción
Description
recovers from crashes during parallel step execution is nondeterministic because it assumes the successful step-a branch has persisted its completion before the failing step-b branch rejects Promise.all. There is no synchronization enforcing that ordering.
Observed failure
AssertionError: expected { a: 'x', b: 'b', attempts: 2 } to deeply equal { a: 'a', b: 'b', attempts: 2 }
- Expected
+ Received
{
- "a": "a",
+ "a": "x",
"attempts": 2,
"b": "b",
}
The test passes most runs but fails intermittently under concurrent/cloud execution.
Race sequence
step-aandstep-bstart concurrently inPromise.all.step-bthrowsSimulated crash.- The workflow is rescheduled and releases worker ownership before
step-adurably records its successful output. - On the second workflow attempt,
step-ais absent from the completed-step cache. - Its callback runs again and deliberately returns
"x"becauseattemptCount > 1.
The assertion expects "a", so correctness currently depends on promise/database scheduling.
Suggested fix
Add explicit synchronization so step-b does not throw until step-a has completed its durable step write. For example, expose a deferred signal resolved after the step-a promise completes, and await it in the step-b branch before throwing. Synchronizing only inside the step-a callback is insufficient because persistence happens after that callback returns.
Alternatively, poll the backend for a completed step-a attempt before allowing step-b to fail. This would preserve the intended assertion: completed parallel work is cached across workflow retry.
Expected behavior
The test should deterministically establish that step-a is completed before triggering the simulated crash, then verify that replay reads "a" from durable history rather than executing its callback again.
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Evaluación
Este issue todavía no se ha evaluado.