microsoft / microsoft/duroxide
Runtime should explicitly handle orphan orchestrator queue messages
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 217
- Forks
- 61
- Avg merge
- 2d 22h
- Merged PRs (30d)
- 3
Description
Problem
When queue messages (e.g., QueueMessage from enqueue_event) arrive in the orchestrator queue before StartOrchestration for a new instance, the runtime's orchestration dispatcher encounters a batch with no instance, no history, and no StartOrchestration/ContinueAsNew message. The current behavior:
fetch_orchestration_itemreturns the batch withorchestration_name="Unknown"- The runtime logs
"completion messages for unstarted instance"and"empty effective batch" - The runtime acks the batch, which permanently deletes the queue rows
- The events are lost forever
This was discovered via the sample_config_hot_reload_persistent_events_fs e2e test, which enqueues events before starting an orchestration.
Current Provider-Side Workaround
Both duroxide-pg and duroxide-pg-opt have implemented a provider-side fix in their fetch_orchestration_item stored procedure:
- Scan ALL messages for
StartOrchestration/ContinueAsNew(not justmessages[0]), matching the SQLite provider'swork_items.iter().find()behavior - If no start item found: release locks and return nothing, leaving messages in the queue until
StartOrchestrationarrives
This works but pushes responsibility to the provider, which:
- Is fragile (providers must each implement this correctly)
- Cannot add a
visible_atdelay to prevent tight re-fetching (any delay risks events being lost if the orchestration completes before the delay expires) - Relies on
LISTEN/NOTIFYfor backpressure to prevent tight-looping
Proposed Runtime-Level Fix
The runtime's orchestration dispatcher should handle this case explicitly:
- When
fetch_orchestration_itemreturns a batch with no instance and noStartOrchestration/ContinueAsNewin the messages, the runtime should abandon the batch (not ack it) - The abandon should use a reasonable delay (e.g., 500ms) so items become available again later
- This keeps the contract simple: providers return whatever is in the queue, and the runtime decides what to do
This would also allow removing the provider-side workarounds.
Provider Validation Test
This issue was only caught by an e2e sample test (sample_config_hot_reload_persistent_events_fs), not by any provider validation test. A dedicated test should be added to duroxide::provider_validation that:
- Enqueues one or more
QueueMessageevents for an instance before callingstart_orchestration - Then starts the orchestration
- Verifies that all pre-enqueued events are delivered to the orchestration (not silently dropped)
This would ensure all provider implementations are validated against this scenario without requiring full e2e tests.
Affected Code
- Runtime:
dispatchers/orchestration.rs- the"completion messages for unstarted instance"code path - Provider trait:
abandon_orchestration_itemis already available for this purpose - Test suite:
duroxide::provider_validation- add orphan message handling test
References
duroxide-pg-optmigration0006_fix_orphan_queue_messages.sqlduroxide-pgmigration0016_fix_orphan_queue_messages.sql- Test:
sample_config_hot_reload_persistent_events_fs
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start in dispatchers/orchestration.rs at the "completion messages for unstarted instance" path, then inspect the existing abandon_orchestration_item provider trait method. Review duroxide::provider_validation and the sample_config_hot_reload_persistent_events_fs test. Done means pre-enqueued QueueMessage events survive until start_orchestration and are delivered without being acknowledged and lost.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- distributed-systems, testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 45/100