apache / apache/beam

Enhance Apache Beam interpreter for Apache Zeppelin

Open
#18,622 0 comments 0 reactions 0 assignees View on GitHub
new feature P3 sdk-ideas sql
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
2d 2h
Merged PRs (30d)
205

Description

Apache Zeppelin includes an integration with Apache Beam: https://zeppelin.apache.org/docs/0.7.0/interpreter/beam.html

How well does this work for interactive exploration? Can this be enhanced to support Beam SQL? What about unbounded data? Let's find out by exploring the existing interpreter and enhancing it particularly for streaming SQL.

This project will require the ability to read, write, and run Java and SQL. You will come out of it with familiarity with two Apache big data projects and lots of ideas!

Imported from Jira [BEAM-3784](https://issues.apache.org/jira/browse/BEAM-3784). Original Jira may contain additional context.
Reported by: kenn.

Contributor guide

Open the contributing guide

Research direction

Start by reading the existing Apache Beam interpreter and the Apache Zeppelin integration documentation linked in the issue. Explore how interactive Java and SQL execution currently works, then investigate the open questions around Beam SQL and unbounded data. Done would require a defined and tested enhancement for streaming SQL support.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, sql
Domain
data-engineering, stream-processing
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.