apache / apache/datafusion-python

Remove pyarrow as required dependency, relying on Arrow PyCapsule Interface

未關閉
#1,227 14 則留言 2 個 reaction 已指派 0 人 在 GitHub 檢視
enhancement
主要語言
Python
星號
604
分支
174
平均合併
1 天 7 小時
30 天內合併 PR
4

描述

**Is your feature request related to a problem or challenge? Please describe what you are trying to do.**

PyArrow is a massive dependency. Unpacked, it tends to be >100MB in size, and, until the latest versions (I think?) also required numpy as its own non-optional dependency.

It's also, in effect the only current dependency
https://github.com/apache/datafusion-python/blob/f0bbad7543717c5f08ba2acb92d42c9d30fd2355/pyproject.toml#L46

It would be great if we could remove it, and that would greatly lessen the minimal environment size for datafusion python.

[Many other Python Arrow libraries](https://github.com/apache/arrow/issues/39195#issuecomment-2245718008) implement the PyCapsule Interface, so the user can use nanoarrow, arro3, Polars, DuckDB, etc, or pyarrow. Whatever is best for them.

**Describe the solution you'd like**

The Arrow PyCapsule Interface is a lightweight, decentralized protocol for sharing Arrow data between Python libraries. We already implement the PyCapsule Interface, so it's just a matter of removing places where we hard-code use of pyarrow.

**Describe alternatives you've considered**

Keep pyarrow dependency.

**Additional context**

貢獻指南

這個儲存庫沒有索引到貢獻指南

研究方向

先檢查第 46 行的 pyproject.toml,並找出 Python bindings 中剩餘的硬編碼 pyarrow 使用處。檢查現有的 Arrow PyCapsule Interface 支援如何在這些路徑中使用。完成的標準是 pyarrow 不再是必要依賴,同時受影響的路徑仍持續支援 Arrow 資料提供者。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
data-engineering
Issue 類型
功能
難度
4/5
預估耗時
3-5 天
活躍度
停滯
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。