oxidecomputer / oxidecomputer/omicron
sagas, especially unwind actions, should wait longer for database claims
Open
Nobody has claimed this yet.
important non-blocker
- Dominant language
- Rust
- Stars
- 572
- Forks
- 97
- Avg merge
- 2d 12h
- Merged PRs (30d)
- 96
Description
I don't have a primary source but I understood from the investigation of #8334 (and oxidecomputer/colo#120) that we have one policy for acquiring qorb claims for the database, which is that we wait up to 30s and then give up with an error.
I'd propose that:
- saga actions should maybe be willing to wait longer?
- saga undo actions should maybe wait indefinitely? I don't think we ever want these to give up for transient issues (which failure to get a claim should always be)
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read the investigation in #8334 and oxidecomputer/colo#120 first, then locate the qorb database-claim acquisition used by saga actions and undo actions. Clarify the intended wait policy for each action type and add coverage showing that transient claim failures receive the agreed behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- backend, databases
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100