Make Spark RDDs readable as PCollections
Open
new feature
P3
runners
spark
- Dominant language
- Java
- Stars
- 8.7k
- Forks
- 4.7k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 196
Description
This could be done by implementing a SparkSource.
Imported from Jira [BEAM-16](https://issues.apache.org/jira/browse/BEAM-16). Original Jira may contain additional context.
Reported by: amitsela.
Contributor guide
Research direction
Start by reviewing the proposed SparkSource and the original BEAM-16 Jira issue for the missing context. Identify the relevant Apache Beam and Spark integration entry points before defining the scope. Done means Spark RDDs can be read as PCollections, with coverage for the supported behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, spark
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100