spark_fdw or provide native Spark integration
Open
tech integration
- Dominant language
- C
- Stars
- 12.8k
- Forks
- 794
- Avg merge
- 2d 14h
- Merged PRs (30d)
- 31
Description
We could consider writing a spark_fdw (foreign data wrapper) to enable querying data in Spark.
Or we could build a tight integration between Spark and PostgreSQL / Citus. In this scenario, Spark manages distributed roll-ups and PostgreSQL acts as the presentation layer.
Contributor guide
Research direction
Start by comparing the two directions described: a spark_fdw for querying Spark data, or tighter Spark integration with PostgreSQL/Citus. Review the existing Spark and PostgreSQL/Citus integration points before choosing an approach. Done means one selected direction is specified and implemented with a clear query or distributed-rollup workflow.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- postgresql, spark
- Domain
- data-engineering, databases, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100