apache / apache/spark

[FEATURE REQUEST] Add ADBC (Arrow Database Connectivity) Data Source

Open
#54,603 9 comments 0 reactions 0 assignees View on GitHub
Dominant language
Scala
Stars
44k
Forks
29.4k
PR merge metrics
No merged PRs in 30d

Description

Add a native ADBC (Arrow Database Connectivity) data source to Spark, similar in spirit to the existing JDBC data source but built on the Arrow-native [ADBC](https://arrow.apache.org/adbc/) API.

ADBC is a database connectivity API standard under the Apache Arrow project. It provides a vendor-neutral, columnar alternative to JDBC/ODBC specifically designed for analytical workloads. ADBC drivers return result sets as streams of Arrow data rather than row-by-row, which eliminates expensive row-to-columnar conversions. Since spark itself is row-based, the effect is not as dramatic, but still noticeable.

Why (now):
- There are mature native drivers for PostgreSQL, SQLite, DuckDB, Flight SQL, Snowflake, BigQuery, MySQL, SQL Server, Databricks and so on. It's also very easy to install (and locate) them on a system with [dbc](https://columnar.tech/dbc/) cli tool.
- There is now good support for invoking ADBC from Java via JNI bindings to the C++ ADBC driver manager (see [blog](https://columnar.tech/blog/adbc-java/)). This makes it practical to integrate ADBC into Spark's JVM-based architecture. Technically drivers can be implemented in java as well, but the quality of java implementations is pretty low, realistically one will almost almost use a native driver.
- ADBC fits well with spark's columnar read support in data source v2. Generating ArrowColumnVectors from adbc is pretty straightforward. It can be a benefit for external spark accelerators like comet and (presumably photon).

I have a proof-of-concept implementation at [spark-adbc](https://github.com/tokoko/spark-adbc) that demonstrates the basic read path and not so scientific benchmarks vs jdbc. I'm willing to incrementally implement ADBC data source support upstream if there's interest from the community.

Contributor guide

Open the contributing guide

Research direction

Review the spark-adbc proof of concept and Spark's existing JDBC data source, then inspect the Data Source V2 and ArrowColumnVectors integration points mentioned in the request. Compare the proof of concept's basic read path and benchmarks with the requested native ADBC support; the issue does not specify a target test or completion scope.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, scala
Domain
data, databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.