apache / apache/datafusion-python

Expose per-file write metadata from DataFrame.write_parquet()

未关闭
#1,637 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
604
派生
174
平均合并
1 天 7 小时
30 天内合并 PR
4

描述

## Is your feature request related to a problem or challenge?

`DataFrame.write_parquet()` currently returns `None`. After writing, there is no way to retrieve per-file metadata (row counts, byte sizes, column statistics) for the files that were produced. This forces consumers that need file-level statistics — such as Apache Iceberg, Delta Lake, and Apache Hudi — to either:

1. Re-read Parquet footers from object storage after writing (extra I/O round-trips)
2. Bypass DataFusion's write pipeline entirely and use PyArrow's `ParquetWriter` with `metadata_collector`

This is a blocker for building a complete DataFusion-based write backend for table formats that require per-file column statistics in their commit metadata (e.g., Iceberg's `DataFile` entries need `column_sizes`, `null_counts`, `lower_bounds`, `upper_bounds`, `split_offsets`).

## Describe the solution you'd like

After [apache/datafusion#23472](https://github.com/apache/datafusion/issues/23472) / [apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) lands in the Rust core, `ParquetSink` will expose a `file_metadata()` method returning per-file path, row count, and byte size. The Python bindings should surface this:

```python
# Option A: write_parquet returns metadata directly
metadata = df.write_parquet("/path/to/output/")
# metadata: list[dict] = [
# {"path": "part-0.parquet", "row_count": 500, "byte_size": 4096},
# {"path": "part-1.parquet", "row_count": 500, "byte_size": 3840},
# ]

# Option B: write_parquet returns a WriteResult object
result = df.write_parquet("/path/to/output/")
result.count # 1000
result.file_metadata # list of per-file metadata dicts
```

At minimum, each file metadata entry should include:
- `path` (str): Object-store path of the written file
- `row_count` (int): Number of rows in this file
- `byte_size` (int): Sum of compressed row group sizes

Optionally (for full table-format integration):
- `metadata` (bytes | None): Serialized Parquet `FileMetaData` (Thrift compact), enabling consumers to extract column statistics without re-reading the file

## Describe alternatives you've considered

- **Return just the count** (status quo): Insufficient for table format integration.
- **Expose via a separate accessor**: e.g. `ctx.last_write_metadata()` — awkward API, not composable.
- **Return raw bytes of the full Parquet footer**: Maximally informative but heavier. A structured dict with optional raw bytes is more ergonomic.

## Additional context

- **Upstream dependency:** [apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) adds `DataSink::file_metadata()` to the Rust core. This issue tracks exposing it through the Python bindings.
- **Motivation:** PyIceberg is building a [pluggable execution backend](https://github.com/apache/iceberg-python/issues/3554) with DataFusion for bounded-memory operations. A DataFusion write backend would enable single-pass Copy-on-Write deletes (read → filter → write entirely in Rust with spill-to-disk), but requires per-file metadata to construct Iceberg `DataFile` commit entries.
- **Related:** #1624 (per-session object store config) is the other piece needed for a complete DataFusion write backend in PyIceberg.

贡献指南

这个仓库没有索引到贡献指南

调研方向

Start by checking whether apache/datafusion#23656 has landed, then read the Python binding for DataFrame.write_parquet() and the linked Rust core issue and pull request. The API shape is still open: the issue suggests returning metadata directly or through a WriteResult. Done means exposing per-file paths, row counts, and byte sizes through the bindings; serialized metadata is optional.

由索引模型根据 Issue 内容生成。

评估

技术栈
python, rust
领域
backend-api-design, data-engineering
Issue 类型
功能
难度
4/5
预计耗时
3-5 天
活跃度
冷清
描述清晰度
基本清楚
新手友好度
48/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。