apache / apache/datafusion-python

Expose per-file write metadata from DataFrame.write_parquet()

Đang mở
#1,637 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Python
Star
604
Fork
174
Merge trung bình
1 ngày 7 giờ
Pull request đã merge (30 ngày)
4

Mô tả

## Is your feature request related to a problem or challenge?

`DataFrame.write_parquet()` currently returns `None`. After writing, there is no way to retrieve per-file metadata (row counts, byte sizes, column statistics) for the files that were produced. This forces consumers that need file-level statistics — such as Apache Iceberg, Delta Lake, and Apache Hudi — to either:

1. Re-read Parquet footers from object storage after writing (extra I/O round-trips)
2. Bypass DataFusion's write pipeline entirely and use PyArrow's `ParquetWriter` with `metadata_collector`

This is a blocker for building a complete DataFusion-based write backend for table formats that require per-file column statistics in their commit metadata (e.g., Iceberg's `DataFile` entries need `column_sizes`, `null_counts`, `lower_bounds`, `upper_bounds`, `split_offsets`).

## Describe the solution you'd like

After [apache/datafusion#23472](https://github.com/apache/datafusion/issues/23472) / [apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) lands in the Rust core, `ParquetSink` will expose a `file_metadata()` method returning per-file path, row count, and byte size. The Python bindings should surface this:

```python
# Option A: write_parquet returns metadata directly
metadata = df.write_parquet("/path/to/output/")
# metadata: list[dict] = [
# {"path": "part-0.parquet", "row_count": 500, "byte_size": 4096},
# {"path": "part-1.parquet", "row_count": 500, "byte_size": 3840},
# ]

# Option B: write_parquet returns a WriteResult object
result = df.write_parquet("/path/to/output/")
result.count # 1000
result.file_metadata # list of per-file metadata dicts
```

At minimum, each file metadata entry should include:
- `path` (str): Object-store path of the written file
- `row_count` (int): Number of rows in this file
- `byte_size` (int): Sum of compressed row group sizes

Optionally (for full table-format integration):
- `metadata` (bytes | None): Serialized Parquet `FileMetaData` (Thrift compact), enabling consumers to extract column statistics without re-reading the file

## Describe alternatives you've considered

- **Return just the count** (status quo): Insufficient for table format integration.
- **Expose via a separate accessor**: e.g. `ctx.last_write_metadata()` — awkward API, not composable.
- **Return raw bytes of the full Parquet footer**: Maximally informative but heavier. A structured dict with optional raw bytes is more ergonomic.

## Additional context

- **Upstream dependency:** [apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) adds `DataSink::file_metadata()` to the Rust core. This issue tracks exposing it through the Python bindings.
- **Motivation:** PyIceberg is building a [pluggable execution backend](https://github.com/apache/iceberg-python/issues/3554) with DataFusion for bounded-memory operations. A DataFusion write backend would enable single-pass Copy-on-Write deletes (read → filter → write entirely in Rust with spill-to-disk), but requires per-file metadata to construct Iceberg `DataFile` commit entries.
- **Related:** #1624 (per-session object store config) is the other piece needed for a complete DataFusion write backend in PyIceberg.

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Start by checking whether apache/datafusion#23656 has landed, then read the Python binding for DataFrame.write_parquet() and the linked Rust core issue and pull request. The API shape is still open: the issue suggests returning metadata directly or through a WriteResult. Done means exposing per-file paths, row counts, and byte sizes through the bindings; serialized metadata is optional.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
python, rust
Lĩnh vực
backend-api-design, data-engineering
Loại issue
Tính năng
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Ít trao đổi
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
48/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.