StackStorm / StackStorm/orquesta

task with join: all starts without waiting for all the previous task completed when there is a loop

Open
#263 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
111
Forks
44
PR merge metrics
No merged PRs in 30d

Description

I made a workflow having 3 parallel tasks each of which randomly fails, along with a task join all of them.
If all of the parallel tasks success, then the workflow will end. Otherwise, the workflow steps back to rerun the parallel tasks. The yaml code is as followed:

version: 1.0

description: loop parallel workflow.

input:
  - x: {}
  - y: {}
  - z: {}

output:
  - data:
      x: <% ctx().x %>
      y: <% ctx().y %>
      z: <% ctx().z %>

tasks:
  entrypoint:
    next: 
      - do: parallel

  parallel:
    action: core.noop
    next:
      - do: random_failure1,random_failure2, random_failure3

  random_failure1:
    action: my_pack.random_failure
    next:
      - publish: x=<% result() %>
        do: count_failure

  random_failure2:
    action: my_pack.random_failure
    next:
      - publish: y=<% result() %>
        do: count_failure

  random_failure3:
    action: my_pack.random_failure
    next:
      - publish: z=<% result() %>
        do: count_failure

  count_failure:
    join: all
    action: my_pack.count_failure
    input:
      - x=<% ctx(x) %>
      - y=<% ctx(y) %>
      - z=<% ctx(z) %> 
    next:
      - when: <% result().failure_count > 0 %>
        do: tasks

In the first trial, the count_failure task starts after all of the 3 parallel task end and gets the context x, y, z updated by parallel tasks. Let s call the context x1, y1, z1.
However, in the second loop, the count_failure task starts after any of the 3 parallel task end. Assuming that random_failure1 ends first and count_failure will start immediately after that, with the context to be x2(updated by random_failure1), y1(not updated), z1(not updated). And when random_failure2 ends, another count_failure task starts with input like x2, y2, z1.

I am expecting the count_failure task also starts starts after all of the 3 parallel task end in the second loop, with input updated as x2, y2, z2.
Is this a bug, or the constraint of the graph based workflow in orquesta?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the supplied YAML workflow with three parallel tasks that can fail and loop back to the parallel branch. Trace how the join-all task is tracked across loop iterations and compare each invocation's context values. Done means determining whether the second iteration's early join is a bug or a documented graph-workflow constraint.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
devops
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.