Datasets should have allowed_states parameter
- Dominant language
- Python
- Stars
- 46.9k
- Forks
- 17.8k
- Avg merge
- 2d 10h
- Merged PRs (30d)
- 483
Description
### Description
Currently, a dataset is updated only when the producing task completes in the success state. I propose adding an `allowed_states` parameter to datasets, which would allow datasets to trigger the downstream consuming DAG even if the producing task is not successful. This would provide more flexibility with dataset scheduling.
Proposed Changes:
- The consuming DAG should include a list of `allowed_states` with the dataset used in the schedule.
- The dataset should be updated once the producing task completes irrespective of the state of the producing task (not only when the producing task succeeds).
- This update should trigger the consuming DAG if the state of the producing task is one of the states included in the `allowed_states` list.
- The `allowed_states` will default to including the `success` state only.
### Use case/motivation
- User wants consuming DAG to be triggered irrespective of the state of the producing task.
- User wants different consuming DAGs to be triggered depending on the state of the producing task.
### Related issues
_No response_
### Are you willing to submit a PR?
- [ ] Yes I am willing to submit a PR!
### Code of Conduct
- [x] I agree to follow this project's [Code of Conduct](https://github.com/apache/airflow/blob/main/CODE_OF_CONDUCT.md)
Contributor guide
Research direction
The issue names no files, tests, or entry points. Start by locating dataset scheduling and task-state handling, then define coverage for the success-only default, dataset updates for completed tasks, and triggering consuming DAGs for configured allowed states.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100