apache / apache/parquet-java

Add way to disables statistics on a per column basis

Offen
#2,521 5 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Component: Java Component: Parquet Priority: Major Type: enhancement
Vorherrschende Sprache
Java
Sterne
3.1k
Forks
1.6k
Ø Merge
3 T. 12 Std.
Gemergte PRs (30 T.)
33

Beschreibung

When you write dataset with BINARY columns that can be fairly large (several Mbs) you can often end with an OutOfMemory error where you either have to:

 

 - Throw more RAM

 - Increase number of output files

 - Play with Block size

 

Using a fork with increased checks frequency for row group size help but it is not enough. (PR: )

 

 

The OutOfMemory error is now caused due to the accumulation of min/max values for those columns for each BlockMetaData.

 

The "parquet.statistics.truncate.length" configuration is of no help because it is applied during the footer serialization whereas the OOM occurs before that.

 

I think it would be nice to have, like for dictionary or bloom filter, a way to disable the statistic on a per-column basis.

 

Could be very useful to lower memory consumption when stats of huge binary column are unnecessary.

 

 

 

**Reporter**: [Anthony Pessy](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=panthony) / @panthony
#### Original Issue Attachments:
- [add_config_to_opt-out_of_a_column's_statistics.patch](https://issues.apache.org/jira/secure/attachment/13011473/add_config_to_opt-out_of_a_column%27s_statistics.patch)
- [NoOpStatistics.java](https://issues.apache.org/jira/secure/attachment/13037932/NoOpStatistics.java)

**Note**: *This issue was originally created as [PARQUET-1911](https://issues.apache.org/jira/browse/PARQUET-1911). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Lesen Sie zunächst die angehängten Dateien NoOpStatistics.java und add_config_to_opt-out_of_a_column's_statistics.patch, und verfolgen Sie anschließend, wie sich Statistiken in BlockMetaData vor der Footer-Serialisierung ansammeln. Als abgeschlossen gilt die Bereitstellung einer Möglichkeit pro Spalte, Statistiken zu deaktivieren, sodass große BINARY-Spalten keine min/max-Werte mehr ansammeln und OutOfMemory-Fehler verursachen.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
java
Bereich
data-engineering
Issue-Typ
Feature
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.