Some liveness checks don't actually check process
- Dominant language
- Python
- Stars
- 46.9k
- Forks
- 17.8k
- Avg merge
- 2d 10h
- Merged PRs (30d)
- 483
Description
### Description
Current liveness check probes use the 'airflow jobs' command which directly queries the backend DB as opposed to actually querying an endpoint or checking the status of the process itself.
e.g. Triggerer liveness probe
```
exec [sh -c CONNECTION_CHECK_MAX_COUNT=0 AIRFLOW__LOGGING__LOGGING_LEVEL=ERROR exec /entrypoint \
airflow jobs check --job-type TriggererJob --hostname $(hostname)] delay=10s timeout=20s period=60s #success=1 #failure=5
```
This command only checks the backend DB to see if there are any jobs. Additionally, the exit code is always 0 regardless of how many jobs there are. Ideally, the liveness check is done by querying some endpoint on the triggerer to see if it's still running.
### Use case/motivation
Would like a liveness check that is more aware of the process rather than the stored state
### Related issues
_No response_
### Are you willing to submit a PR?
- [ ] Yes I am willing to submit a PR!
### Code of Conduct
- [X] I agree to follow this project's [Code of Conduct](https://github.com/apache/airflow/blob/main/CODE_OF_CONDUCT.md)
Contributor guide
Research direction
Start by tracing the Triggerer liveness probe and the `airflow jobs check --job-type TriggererJob` entry point shown in the issue. Confirm how the command determines its exit status and whether the probe observes the process or only backend database state. Done means the affected liveness check reflects actual process health and reports failure when that process is not healthy.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- observability
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100