apache / apache/pinot

Apache Beam IO Connector

Open
#7,147 1 comment 1 reaction 0 assignees View on GitHub
feature ingestion
Dominant language
Java
Stars
6.1k
Forks
1.5k
Avg merge
2d 3h
Merged PRs (30d)
195

Description

Hi folks 👋

I noted that there's an Apache Spark integration[1] for creating segments and batch-ingesting data. I'd be keen to explore creating a similar IO connector for Apache Beam[2], and I would be interested in providing support both in terms of brainstorming and/or implementation.

I'm opening this issue to better understand the community support of and implications of such an integration, what related work might already be underway, and how generalizable the core components of `org.apache.pinot.plugin.ingestion.batch.spark` might be.

[1] https://docs.pinot.apache.org/basics/data-import/batch-ingestion/spark
[2] https://beam.apache.org/documentation/io/developing-io-overview/

Contributor guide

Open the contributing guide

Research direction

Review the Apache Spark batch-ingestion documentation and the org.apache.pinot.plugin.ingestion.batch.spark package, then compare them with Apache Beam's IO connector guidance. Clarify whether a Beam connector is already underway, which Spark components are generalizable, and what scope and acceptance criteria an implementation would need.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.