huggingface / huggingface/datatrove
Optimize parquet output for remote reading
- Dominant language
- Python
- Stars
- 3.3k
- Forks
- 302
- Avg merge
- 2h 18m
- Merged PRs (30d)
- 2
Description
TLDR: the primary pain point here is huge (in terms of total uncompressed byte size) row groups - writing the PageIndex OR reducing row group sizes, perhaps both, would help a lot.
Basically, the defaults in pyarrow (and most parquet implementations) for row group sizes (1 million rows per row group) are predicated on assumptions about what a typical parquet file looks like (lots of numerics, booleans, relatively short strings amenable to dictionary, RLE and delta encoding); wide text datasets are very much not typical, and default row group sizes get you ~2GB per row group (and nearly 4GB uncompressed, just for the text column).
The simplest thing to do would be to default to 100k for the row_group_size parameter - more or less the inflection point of [this](https://duckdb.org/docs/guides/performance/file_formats.html#parquet-file-sizes) benchmark by DuckDB (size overhead is about 0.5%).
Setting `write_page_index` to true should help a great deal (arguably much more than smaller row groups), as readers can use that to refine reads to individual data pages (not unusual for point lookups to hit 0.1% of a file).
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.