apache / apache/arrow

[C++][Parquet][CI] Improve Parquet fuzzing seed corpus

Open
#43,709 0 comments 0 reactions 1 assignee Claimed by @pitrou View on GitHub
Component: C++ Component: Continuous Integration Component: Parquet Type: enhancement
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the enhancement requested

Currently, for our Parquet fuzzing seed corpus, we generate a grand total of 1 file here:
https://github.com/apache/arrow/blob/fb202ee66d73572f46035c5b2f21ac22f74ba951/cpp/src/parquet/arrow/generate_fuzz_corpus.cc

We should probably generate more files (and/or more batch columns) and/or enable more features:
* vary data page version
* vary compression codec
* vary encodings (e.g. delta binary, byte stream split...)
* enable page checksums (and verify them on reading: that's actually a bad idea as it would prevent exercising the actual decoding most of the time)
* enable statistics (and load them on reading)
* enable page indices
* enable bloom filters once https://github.com/apache/arrow/pull/37400 is merged

We should also add more datatypes, at least Boolean and FixedSizeBinary, possibly also Decimal128 and Decimal256.

### Component(s)

C++, Continuous Integration, Parquet

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.