citusdata / citusdata/citus

spark_fdw or provide native Spark integration

Open
#47 6 comments 0 reactions 0 assignees View on GitHub
tech integration
Dominant language
C
Stars
12.8k
Forks
794
Avg merge
2d 14h
Merged PRs (30d)
31

Description

We could consider writing a spark_fdw (foreign data wrapper) to enable querying data in Spark.

Or we could build a tight integration between Spark and PostgreSQL / Citus. In this scenario, Spark manages distributed roll-ups and PostgreSQL acts as the presentation layer.

Contributor guide

Open the contributing guide

Research direction

Start by comparing the two directions described: a spark_fdw for querying Spark data, or tighter Spark integration with PostgreSQL/Citus. Review the existing Spark and PostgreSQL/Citus integration points before choosing an approach. Done means one selected direction is specified and implemented with a clear query or distributed-rollup workflow.

Written by the indexing model from the issue text.

Assessment

Tech stack
postgresql, spark
Domain
data-engineering, databases, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.