oxidecomputer / oxidecomputer/omicron

sagas, especially unwind actions, should wait longer for database claims

Open
#8,423 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

important non-blocker
Dominant language
Rust
Stars
572
Forks
97
Avg merge
2d 12h
Merged PRs (30d)
96

Description

I don't have a primary source but I understood from the investigation of #8334 (and oxidecomputer/colo#120) that we have one policy for acquiring qorb claims for the database, which is that we wait up to 30s and then give up with an error.

I'd propose that:

  • saga actions should maybe be willing to wait longer?
  • saga undo actions should maybe wait indefinitely? I don't think we ever want these to give up for transient issues (which failure to get a claim should always be)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Read the investigation in #8334 and oxidecomputer/colo#120 first, then locate the qorb database-claim acquisition used by saga actions and undo actions. Clarify the intended wait policy for each action type and add coverage showing that transient claim failures receive the agreed behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
backend, databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.