apache / apache/datafusion-python

Decode Python UDFs opaquely so a scheduler needs no Python interpreter

未关闭
#1,705 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
enhancement rust
主要语言
Python
星标
604
派生
174
平均合并
1 天 7 小时
30 天内合并 PR
4

描述

**Is your feature request related to a problem or challenge? Please describe what you are trying to do.**

In a distributed setup the scheduler plans a query and hands stages to executors; only the executors ever call a Python UDF. Decoding an inlined Python UDF unpickles the function, which requires a Python interpreter and every module the function closes over to be importable. Doing that on the scheduler costs work nobody needs and forces the scheduler image to carry Python and the full dependency set of user code it will never run. Raised in https://github.com/apache/datafusion-python/pull/1678#pullrequestreview-5100366976.

**Describe the solution you'd like**

An opaque decode path: a `ScalarUDFImpl` that holds the still-pickled blob rather than a live Python object, and a codec that produces it. A scheduler installs that codec, decodes a plan into something it can inspect, route, and re-encode, and never touches cloudpickle. The executor installs the ordinary codec and unpickles as it does now.

The wire format already allows this. An inlined UDF payload is `DFPYUDF` followed by a version byte and the cloudpickle blob (`crates/core/src/codec.rs`), so an opaque holder can carry those bytes verbatim and no format change is needed.

The part that needs design is re-encoding. A scheduler that forwards a stage has to emit the blob byte-identically, so the executor sees exactly what the client wrote. That also raises what such a UDF should report for the things DataFusion asks of a `ScalarUDFImpl` during planning — name, signature, and return type are all recoverable from the payload without unpickling, since they are stored alongside the function, but `invoke` has to be an error rather than a surprise.

**Describe alternatives you've considered**

Encoding Python UDFs by name only and registering them on every node. Already supported and appropriate when the function is available everywhere; it does not cover the case inlining exists for, which is a function the receiving process does not have.

Having the scheduler unpickle and immediately drop the object. Keeps the code simple, and still requires Python plus all user dependencies on the scheduler, which is the actual cost being avoided.

**Additional context**

Follow-up from #1678, which made extension codecs compose so a setup like this can install a scheduler-side codec alongside others. Likely also depends on #1703, gating `pyo3/extension-module`, if the consumer is a Rust crate rather than a Python process.

贡献指南

这个仓库没有索引到贡献指南

调研方向

首先阅读 crates/core/src/codec.rs 以及 #1678 中的扩展 codec 工作,以了解 DFPYUDF wire format 和 codec 组合方式。围绕现有 payload 设计 scheduler 侧的 ScalarUDFImpl 和 codec,在重新编码时保留 cloudpickle 字节,同时暴露可恢复的元数据,并让 invoke 返回错误。检查 #1703 中的依赖上下文。

由索引模型根据 Issue 内容生成。

评估

技术栈
python, rust
领域
backend, distributed-systems
Issue 类型
功能
难度
5/5
预计耗时
一周以上
活跃度
活跃
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。