apache / apache/airflow

Airflow 3.2.0 scheduler/triggerer deadlock on task_instance due to concurrent updates of deferrable tasks

Open
#65,818 6 comments 4 reactions 0 assignees View on GitHub
area:core area:scheduler area:Triggerer kind:bug needs-triage priority:critical
Dominant language
Python
Stars
46.9k
Forks
17.8k
Avg merge
2d 9h
Merged PRs (30d)
472

Description

### Under which category would you file this issue?

Airflow Core

### Apache Airflow version

3.2.0

### What happened and how to reproduce it?

# Description

After upgrading from Airflow 3.1.7 → 3.2.0, we are consistently observing MySQL deadlocks between the scheduler and triggerer when processing deferrable tasks.

This did not occur in 3.1.7 under the same workload.

# Environment
- Airflow version: 3.2.0
- Previous version (no issue): 3.1.7
- Executor: CeleryExecutor
- DB: MySQL
- Scheduler replicas: 3
- Triggerer: 1 instance
- Workload: heavy use of deferrable operators (sensors / async tasks)
# Symptoms
- Scheduler/Trigger crashes or restarts due to DB deadlocks
- Deadlocks consistently involve task_instance table
- System becomes unstable under load

Example deadlock pattern:
```
UPDATE task_instance
SET updated_at=..., trigger_id=NULL
WHERE task_instance.state != 'deferred'
AND task_instance.trigger_id IS NOT NULL
```

conflicting with:
```
UPDATE task_instance
SET state='scheduled', trigger_id=NULL,
next_method='__fail__', next_kwargs=...
WHERE task_instance.state = 'deferred'
AND task_instance.trigger_timeout < now()
```
# Root Cause Analysis
## Key observation

In Airflow 3.2.0, both scheduler and triggerer mutate task_instance rows for deferrable tasks:

## Triggerer (set-based update)
- Performs bulk UPDATE on deferred tasks that timeout
- Updates:
- state
- trigger_id
- next_method
- next_kwargs
## Scheduler (callback-driven updates)
- Processes executor callbacks via:
```
callback = session.get(Callback, callback_id)
callback.run(session=session)
```
- Inside callback.run():
- Loads TaskInstance
- Mutates:
- state
- trigger_id
- other fields
## Result

Two independent writers:

- Triggerer → bulk UPDATE (set-based)
- Scheduler → row-by-row ORM UPDATE

Both target overlapping `task_instance` rows.

## Why this causes deadlocks
- Both queries scan overlapping row sets (even if predicates are logically disjoint)
- Lock acquisition order differs:
- Triggerer: index scan order
- Scheduler: callback / primary key order
- With multiple scheduler replicas, contention increases significantly

Typical pattern:
```
Scheduler: locks row A → waits for row B
Triggerer: locks row B → waits for row A
→ DEADLOCK
```

### What you think should happen instead?

1. Avoid concurrent writes:
- Scheduler should not mutate task_instance fields owned by triggerer
2. Enforce consistent ordering:
- Ensure both components lock rows in deterministic order
3. Batch updates:
- Avoid large scans or uncontrolled ORM flushes
4. Ownership separation:
- Triggerer handles deferred lifecycle exclusively
- Scheduler only consumes results

### Operating System

Ubuntu 22.04.5 LTS

### Deployment

None

### Apache Airflow Provider(s)

_No response_

### Versions of Apache Airflow Providers

_No response_

### Official Helm Chart version

Not Applicable

### Kubernetes Version

_No response_

### Helm Chart configuration

_No response_

### Docker Image customizations

_No response_

### Anything else?

_No response_

### Are you willing to submit PR?

- [ ] Yes I am willing to submit a PR!

### Code of Conduct

- [x] I agree to follow this project's [Code of Conduct](https://github.com/apache/airflow/blob/main/CODE_OF_CONDUCT.md)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.