apache / apache/datafusion-python
Expose per-file write metadata from DataFrame.write_parquet()
- Ngôn ngữ chính
- Python
- Star
- 604
- Fork
- 174
- Merge trung bình
- 1 ngày 7 giờ
- Pull request đã merge (30 ngày)
- 4
Mô tả
## Is your feature request related to a problem or challenge?
`DataFrame.write_parquet()` currently returns `None`. After writing, there is no way to retrieve per-file metadata (row counts, byte sizes, column statistics) for the files that were produced. This forces consumers that need file-level statistics — such as Apache Iceberg, Delta Lake, and Apache Hudi — to either:
1. Re-read Parquet footers from object storage after writing (extra I/O round-trips)
2. Bypass DataFusion's write pipeline entirely and use PyArrow's `ParquetWriter` with `metadata_collector`
This is a blocker for building a complete DataFusion-based write backend for table formats that require per-file column statistics in their commit metadata (e.g., Iceberg's `DataFile` entries need `column_sizes`, `null_counts`, `lower_bounds`, `upper_bounds`, `split_offsets`).
## Describe the solution you'd like
After [apache/datafusion#23472](https://github.com/apache/datafusion/issues/23472) / [apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) lands in the Rust core, `ParquetSink` will expose a `file_metadata()` method returning per-file path, row count, and byte size. The Python bindings should surface this:
```python
# Option A: write_parquet returns metadata directly
metadata = df.write_parquet("/path/to/output/")
# metadata: list[dict] = [
# {"path": "part-0.parquet", "row_count": 500, "byte_size": 4096},
# {"path": "part-1.parquet", "row_count": 500, "byte_size": 3840},
# ]
# Option B: write_parquet returns a WriteResult object
result = df.write_parquet("/path/to/output/")
result.count # 1000
result.file_metadata # list of per-file metadata dicts
```
At minimum, each file metadata entry should include:
- `path` (str): Object-store path of the written file
- `row_count` (int): Number of rows in this file
- `byte_size` (int): Sum of compressed row group sizes
Optionally (for full table-format integration):
- `metadata` (bytes | None): Serialized Parquet `FileMetaData` (Thrift compact), enabling consumers to extract column statistics without re-reading the file
## Describe alternatives you've considered
- **Return just the count** (status quo): Insufficient for table format integration.
- **Expose via a separate accessor**: e.g. `ctx.last_write_metadata()` — awkward API, not composable.
- **Return raw bytes of the full Parquet footer**: Maximally informative but heavier. A structured dict with optional raw bytes is more ergonomic.
## Additional context
- **Upstream dependency:** [apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) adds `DataSink::file_metadata()` to the Rust core. This issue tracks exposing it through the Python bindings.
- **Motivation:** PyIceberg is building a [pluggable execution backend](https://github.com/apache/iceberg-python/issues/3554) with DataFusion for bounded-memory operations. A DataFusion write backend would enable single-pass Copy-on-Write deletes (read → filter → write entirely in Rust with spill-to-disk), but requires per-file metadata to construct Iceberg `DataFile` commit entries.
- **Related:** #1624 (per-session object store config) is the other piece needed for a complete DataFusion write backend in PyIceberg.
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Hướng nghiên cứu
Start by checking whether apache/datafusion#23656 has landed, then read the Python binding for DataFrame.write_parquet() and the linked Rust core issue and pull request. The API shape is still open: the issue suggests returning metadata directly or through a WriteResult. Done means exposing per-file paths, row counts, and byte sizes through the bindings; serialized metadata is optional.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Đánh giá
- Công nghệ
- python, rust
- Lĩnh vực
- backend-api-design, data-engineering
- Loại issue
- Tính năng
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức độ hoạt động
- Ít trao đổi
- Độ rõ ràng
- Khá rõ ràng
- Mức phù hợp với người mới
- 48/100