Apache Beam IO Connector
- Dominant language
- Java
- Stars
- 6.1k
- Forks
- 1.5k
- Avg merge
- 2d 3h
- Merged PRs (30d)
- 195
Description
Hi folks 👋
I noted that there's an Apache Spark integration[1] for creating segments and batch-ingesting data. I'd be keen to explore creating a similar IO connector for Apache Beam[2], and I would be interested in providing support both in terms of brainstorming and/or implementation.
I'm opening this issue to better understand the community support of and implications of such an integration, what related work might already be underway, and how generalizable the core components of `org.apache.pinot.plugin.ingestion.batch.spark` might be.
[1] https://docs.pinot.apache.org/basics/data-import/batch-ingestion/spark
[2] https://beam.apache.org/documentation/io/developing-io-overview/
Contributor guide
Research direction
Review the Apache Spark batch-ingestion documentation and the org.apache.pinot.plugin.ingestion.batch.spark package, then compare them with Apache Beam's IO connector guidance. Clarify whether a Beam connector is already underway, which Spark components are generalizable, and what scope and acceptance criteria an implementation would need.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, spark
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100