conductor-oss / conductor-oss/conductor

Two Workflow Instances Open on Failure

Open
#284 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Java
Stars
32.2k
Forks
1k
Avg merge
2d 5h
Merged PRs (30d)
41

Description

**Description:**
Since upgrading to **Conductor 3.16.0**, we have encountered unusual behavior in one of our workflows. The workflow is defined as follows:
![conductor issue](https://github.com/user-attachments/assets/49faca7d-8c35-4d36-a56f-c080447cba7c)

When the **WAIT EVENT** receives a message, the workflow proceeds to the **TERMINATE TASK**. However, occasionally we observe that two failure workflows are opened. These failure workflows are nearly identical, but one shows an ownerApp as "conductor," while the other has an empty ownerApp.

Main workflow:

{

  "ownerApp": "",

  "createTime": 1728474259443,

  "updateTime": 1728477038331,

  "status": "FAILED",

  "endTime": 1728477038331,

  "workflowId": "179dd481-8e4b-4e9d-905d-8f37f9b7c577",

  "tasks": […]

}

Output:

{

  "output": "",

  "conductor.failure_workflow": "5974405e-e4b6-4924-b0bf-fbcec3827e2b"

}

failure workflow 1:

{

  "**ownerApp**": "conductor",

  "**createTime**": 1728477038**313**,

  "updateTime": 1728477039058,

  "status": "COMPLETED",

  "**endTime**": 1728477039**058**,

  "workflowId": "5974405e-e4b6-4924-b0bf-fbcec3827e2b",

  "tasks": […]

}

failure workflow 2:

{

  "**ownerApp**": "",

  "**createTime**": 1728477038**229**,

  "updateTime": 1728477039145,

  "status": "COMPLETED",

  "**endTime**": 1728477039**145**,

  "workflowId": "c02b2cb8-d6c4-4aaa-bc1a-3c04a1585d80",

  "tasks": […]

}

From the main workflow output, the failure workflow ID corresponds to the one with the ownerApp set to "conductor." The timestamps show that the two workflows are opened just a few milliseconds apart.

Here are the relevant logs for further insight:

![image (1)](https://github.com/user-attachments/assets/2d0a7b45-600f-4495-8f05-c2eea9d5df68)

Based on these logs, we suspect that this behavior may be caused by race conditions on the workflow's status. It seems related to the **sweeper thread** triggering an action while the event is already being processed by the main flow.

**Expected Behavior:** Only one failure workflow instance should be opened when the workflow fails.

**Potential Cause:** The issue appears to be caused by a race condition in the **decider queue**, specifically around status updates when the workflow progresses from the WAIT EVENT to the TERMINATE TASK. The sweeper thread may be triggering actions prematurely, while the event processing is still ongoing in the main workflow flow.

Contributor guide

Open the contributing guide

Research direction

Start by tracing the decider queue and sweeper thread while following the WAIT EVENT to TERMINATE TASK transition. Use the provided workflow records and logs to compare the near-simultaneous status updates and failure-workflow creation. Done means reproducing the race and ensuring a failed workflow opens only one failure workflow instance.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.