apache / apache/datafusion-java
Spark DataSource backed by a DataFusion TableProvider over ADBC
- 主要語言
- Java
- 星號
- 32
- 分支
- 12
- PR 合併指標
- 30 天內沒有已合併 PR
描述
**Is your feature request related to a problem or challenge?**
Spark users want to read data from a DataFusion `TableProvider` as a native Spark `DataSourceV2`. Today there is no first-class path; options are either a bespoke per-operation JNI surface (more native surface to maintain) or copying data out of process.
**Describe the solution you'd like**
A Spark `DataSourceV2` connector that places the native boundary at a **standard ADBC driver**. Spark talks to the upstream arrow-adbc Java driver manager (`adbc-core` + `adbc-driver-jni`), which loads a native DataFusion ADBC cdylib and returns arrow-java `ArrowReader`s consumed zero-copy as `ArrowColumnVector`s on the cluster-provided Arrow. This reuses the upstream ADBC bindings rather than reproducing them.
Scope:
- `adbc-datafusion` format registered as a `DataSourceV2`; schema probed on the driver.
- Projection / filter / limit pushdown via Substrait, with a SQL fallback.
- Multi-partition reads (`executePartitioned` / `readPartition`) and a `target_partitions` option.
- Per-executor connection pool to amortize driver/database setup across task slots.
- An example DataFusion ADBC driver cdylib plus end-to-end (PySpark) coverage.
**Describe alternatives you've considered**
A plain-C scan ABI + hand-written JNI shim (discussed on #103 / #104). The ADBC approach reuses standard, separately-reviewed bindings and a stable driver contract instead.
**Additional context**
Implemented in #111.
貢獻指南
研究方向
先從 #111 中引用的實作開始,然後將其與此 issue 宣告的範圍進行比較:adbc-datafusion DataSourceV2、pushdowns、分割讀取、executor connection pooling 和 PySpark 覆蓋。執行 issue 中提到的端到端覆蓋,並驗證列出的每項能力都已得到體現且正常運作。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- java, python, spark
- 領域
- backend, data-engineering, distributed-systems
- Issue 類型
- 功能
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 停滯
- 描述清晰度
- 描述清楚
- 新手友好度
- 25/100