Azure / Azure/durabletask

Azure Storage backend: control queue partition left unowned for hours/days after lease expires

Open
#1,389 1 comment 0 reactions 1 assignee Claimed by @nytian View on GitHub
Dominant language
C#
Stars
1.7k
Forks
335
Avg merge
2d 23h
Merged PRs (30d)
6

Description

**Environment**
Microsoft.Azure.DurableTask.AzureStorage: 2.9.1
Microsoft.Azure.DurableTask.Core: 3.9.0
Runtime: .NET 8
Host: Kubernetes, self-hosted TaskHubWorker (not Azure Functions)
Partition manager: Table partition manager
PartitionCount: 16
Worker replicas: 25
UseAppLease: default (true)
Non-default settings: TaskHubName, PartitionCount, StorageAccountClientProvider, LoggerFactory only

**Summary**
A single control queue partition stopped being polled by any worker for hours/days. Messages continued to be enqueued to it, but nothing dequeued them until an unrelated worker eventually acquired the partition, at which point many stranded messages were drained at once (many were  ExecutionStarted). Maximum observed message delay was 3 days for some cases.
No worker process crashed, no node was lost, and the previous owner never released the partition.
For example, sharing one request timeline for one partition:

Timeline (one partition, referred to below as control-NN)

T+00:07 Lease thrash: acquire → LeaseLost → DrainTablePartitionAsync → re-acquire → DropLostControlQueue. Occurs
T+00:21 twice.
T+02:26 Last message consumed from this partition. No lease release is logged.
T+02:26 16 messages enqueued. Zero dequeues.
T+17:23
T+17:23 A worker acquires the partition and drains 17 stranded messages.

**Evidence the partition was unowned**
No new messages were found backing off is emitted by the control queue reader loop, so its rate tracks fetch attempts ~1:1 whenever a partition is owned

Hour │ Fetches │ Backoff log lines │
01 │ 62 │ 61 │
02 │ 2 │ 2 │
04–15 │ 0 │ 0 │
18 │ 18 │ 18 │

Zero backoff lines during the outage rules out "owned but concurrency-starved" no worker was running a reader loop against this partition at all.
With  LeaseAcquireInterval = 10s and LeaseInterval = 30s, expected reclaim after an owner stops renewing is ~10–30 seconds.

We were expecting to behave the partition normally, but it remained unowned for ~15.3 hours while the other workers continued polling the remaining partitions on the same task hub normally.

**Impact**
Orchestrations whose  ExecutionStarted  message landed on the affected partition remained in Pending indefinitely, then began executing 10–15 hours later. By that point the external resources they operate on had already been cleaned up by a separate lifecycle process, so the orchestrations failed with misleading downstream errors, which made the underlying cause difficult to attribute.
Activities were unaffected, since the work-item queue is unpartitioned and polled by all workers. Only orchestration-level messages (which ride the partitioned control queues) were impacted.

**Recurrence**
Occurring regularly for with about

**Questions**
1. Is this a known defect in the table partition manager in 2.9.x, and is it addressed in a later release?
2. Can DropLostControlQueue/DrainTablePartitionAsync  following a  LeaseLost  leave a partition in a state where no worker re-acquires it, or where the owner record isn't cleaned up?
3. Is running more workers (25) than partitions (16) known to aggravate lease contention in this path? We are considering to reduced replicas to match PartitionCount in our next change.
4. Is UseAppLease = true appropriate for self-hosted (non-Functions) deployments?
5. Any recommended mitigation or detection while a fix is pending?

**Mitigation / detection we are considering in next change**
• Restructured orchestrations to prefer activities over sub-orchestrations, so more work rides the unpartitioned work-item queue
• Reduced worker replicas to match  PartitionCount 
• Added a monitor that alerts when instances remain in Pending beyond a threshold, since a never-dispatched orchestration emits no telemetry of its own

I can provide additional detail if required.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.