apache / apache/datafusion-python
Expose per-file write metadata from DataFrame.write_parquet()
- Dominant language
- Python
- Stars
- 604
- Forks
- 174
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 4
Description
## Is your feature request related to a problem or challenge?
`DataFrame.write_parquet()` currently returns `None`. After writing, there is no way to retrieve per-file metadata (row counts, byte sizes, column statistics) for the files that were produced. This forces consumers that need file-level statistics — such as Apache Iceberg, Delta Lake, and Apache Hudi — to either:
1. Re-read Parquet footers from object storage after writing (extra I/O round-trips)
2. Bypass DataFusion's write pipeline entirely and use PyArrow's `ParquetWriter` with `metadata_collector`
This is a blocker for building a complete DataFusion-based write backend for table formats that require per-file column statistics in their commit metadata (e.g., Iceberg's `DataFile` entries need `column_sizes`, `null_counts`, `lower_bounds`, `upper_bounds`, `split_offsets`).
## Describe the solution you'd like
After [apache/datafusion#23472](https://github.com/apache/datafusion/issues/23472) / [apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) lands in the Rust core, `ParquetSink` will expose a `file_metadata()` method returning per-file path, row count, and byte size. The Python bindings should surface this:
```python
# Option A: write_parquet returns metadata directly
metadata = df.write_parquet("/path/to/output/")
# metadata: list[dict] = [
# {"path": "part-0.parquet", "row_count": 500, "byte_size": 4096},
# {"path": "part-1.parquet", "row_count": 500, "byte_size": 3840},
# ]
# Option B: write_parquet returns a WriteResult object
result = df.write_parquet("/path/to/output/")
result.count # 1000
result.file_metadata # list of per-file metadata dicts
```
At minimum, each file metadata entry should include:
- `path` (str): Object-store path of the written file
- `row_count` (int): Number of rows in this file
- `byte_size` (int): Sum of compressed row group sizes
Optionally (for full table-format integration):
- `metadata` (bytes | None): Serialized Parquet `FileMetaData` (Thrift compact), enabling consumers to extract column statistics without re-reading the file
## Describe alternatives you've considered
- **Return just the count** (status quo): Insufficient for table format integration.
- **Expose via a separate accessor**: e.g. `ctx.last_write_metadata()` — awkward API, not composable.
- **Return raw bytes of the full Parquet footer**: Maximally informative but heavier. A structured dict with optional raw bytes is more ergonomic.
## Additional context
- **Upstream dependency:** [apache/datafusion#23656](https://github.com/apache/datafusion/pull/23656) adds `DataSink::file_metadata()` to the Rust core. This issue tracks exposing it through the Python bindings.
- **Motivation:** PyIceberg is building a [pluggable execution backend](https://github.com/apache/iceberg-python/issues/3554) with DataFusion for bounded-memory operations. A DataFusion write backend would enable single-pass Copy-on-Write deletes (read → filter → write entirely in Rust with spill-to-disk), but requires per-file metadata to construct Iceberg `DataFile` commit entries.
- **Related:** #1624 (per-session object store config) is the other piece needed for a complete DataFusion write backend in PyIceberg.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by checking whether apache/datafusion#23656 has landed, then read the Python binding for DataFrame.write_parquet() and the linked Rust core issue and pull request. The API shape is still open: the issue suggests returning metadata directly or through a WriteResult. Done means exposing per-file paths, row counts, and byte sizes through the bindings; serialized metadata is optional.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, rust
- Domain
- backend-api-design, data-engineering
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100