apache / apache/beam

[Task]: Add UseDataStreamForBatch option to the Flink runner.

Open
#25,740 0 comments 0 reactions 0 assignees View on GitHub
awaiting triage flink P3 task
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
1d 20h
Merged PRs (30d)
196

Description

### What needs to happen?

This is the second task for migrating Flink batch job execution from DataSet to DataStream API. We will add a new option of `UseDataStreamForBatch` to the Flink runner. When it is set true, the Flink runner will use DataStream to execute the batch jobs, otherwise the existing DataSet API will be used.

### Issue Priority

Priority: 2 (default / most normal work should be filed as P2)

### Issue Components

- [ ] Component: Python SDK
- [ ] Component: Java SDK
- [ ] Component: Go SDK
- [ ] Component: Typescript SDK
- [ ] Component: IO connector
- [ ] Component: Beam examples
- [ ] Component: Beam playground
- [ ] Component: Beam katas
- [ ] Component: Website
- [ ] Component: Spark Runner
- [X] Component: Flink Runner
- [ ] Component: Samza Runner
- [ ] Component: Twister2 Runner
- [ ] Component: Hazelcast Jet Runner
- [ ] Component: Google Cloud Dataflow Runner

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.