apache / apache/parquet-java

Unnecessary getFileStatus() calls on all part-files in ParquetInputFormat.getSplits

Offen
#1,417 2 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Component: Java Component: Parquet Priority: Major Type: bug
Vorherrschende Sprache
Java
Sterne
3.1k
Forks
1.6k
Ø Merge
3 T. 12 Std.
Gemergte PRs (30 T.)
33

Beschreibung

When testing Spark SQL Parquet support, we found that accessing large Parquet files located in S3 can be very slow. To be more specific, we have a S3 Parquet file with over 3,000 part-files, calling `ParquetInputFormat.getSplits` on it takes several minutes. (We were accessing this file from our office network rather than AWS.)

After some investigation, we found that `ParquetInputFormat.getSplits` is trying to call `getFileStatus()` on all part-files one by one sequentially ([here](https://github.com/apache/incubator-parquet-mr/blob/parquet-1.5.0/parquet-hadoop/src/main/java/parquet/hadoop/ParquetInputFormat.java#L370)). And in the case of S3, each `getFileStatus()` call issues an HTTP request and wait for the reply in a blocking manner, which is considerably expensive.

Actually all these `FileStatus` objects have already been fetched when footers are retrieved ([here](https://github.com/apache/incubator-parquet-mr/blob/parquet-1.5.0/parquet-hadoop/src/main/java/parquet/hadoop/ParquetInputFormat.java#L443)). Caching these `FileStatus` objects can greatly improve our S3 case (reduced from over 5 minutes to about 1.4 minutes).

Will submit a PR for this issue soon.

**Reporter**: [Cheng Lian](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=lian+cheng) / @liancheng
#### Related issues:
- [Use LRU caching for footers in ParquetInputFormat.](https://github.com/apache/parquet-java/issues/1394) (relates to)
- [Cleanup FilteringParquetRowInputFormat](https://issues.apache.org/jira/browse/SPARK-2551) (is related to)
- [Reading Parquet InputSplits dominates query execution time when reading off S3](https://issues.apache.org/jira/browse/SPARK-2119) (is related to)
- [Improve Parquet IO Performance within cloud datalakes](https://github.com/apache/parquet-java/issues/2912) (is depended upon by)

**Note**: *This issue was originally created as [PARQUET-16](https://issues.apache.org/jira/browse/PARQUET-16). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Beginne in parquet-hadoop/src/main/java/parquet/hadoop/ParquetInputFormat.java bei getSplits und dem im Issue verknüpften Pfad zum Abrufen der Footer. Verfolge, wie FileStatus-Objekte abgerufen werden, und überprüfe anschließend, dass getSplits sie nicht mehr sequenziell für jede part-file abruft. Vergleiche außerdem die im Report beschriebene Zeitmessung für S3 oder große Dateien.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
java
Bereich
performance
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.