Update Python SDK to construct Dataflow job requests from Beam runner API protos
- Dominant language
- Java
- Stars
- 8.7k
- Forks
- 4.7k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 196
Description
Currently, portable runners are expected to do following when constructing a runner specific job.
SDK specific job graph -\> Beam runner API proto -\> Runner specific job request
Portable Spark and Flink follow this model.
Dataflow does following.
SDK specific job graph -\> Runner specific job request
Beam runner API proto -\> Upload to GCS -\> Download at workers
We should update Dataflow to follow the prior path which is expected to be followed by all portable runners.
This will simplify the cross-language transforms job construction logic for Dataflow.
We can probably start this by just implementing this for Python SDK for portions of pipeline received by expanding external transforms.
cc: [~lcwik] [~robertwb]
Imported from Jira [BEAM-10012](https://issues.apache.org/jira/browse/BEAM-10012). Original Jira may contain additional context.
Reported by: chamikara.
Contributor guide
Assessment
This issue has not been assessed yet.