apache / apache/parquet-java

Cache ParsedVersion in FileMetaData to eliminate redundant created_by parsing

Offen
#3,696 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Java
Sterne
3.1k
Forks
1.6k
Ø Merge
3 T. 12 Std.
Gemergte PRs (30 T.)
33

Beschreibung

## Summary

`VersionParser.parse(createdBy)` is called from 7 distinct production sites, all parsing the same constant string from `FileMetaData.getCreatedBy()`. Since `FileMetaData` is constructed once per file and already stores the `createdBy` string, it is the natural place to parse once and cache.

## Problem

The `created_by` string is re-parsed into a `ParsedVersion` at every call site independently:

| # | Call site | Frequency per file | When |
|---|-----------|-------------------|------|
| 1 | `CorruptStatistics.shouldIgnoreStatistics` via `buildColumnChunkMetaData` | R × C | Footer decode |
| 2 | `CorruptStatistics.shouldIgnoreStatistics` via `ParquetFileReader.readAllPages` | Pages per column chunk | Page read |
| 3 | `CorruptStatistics.shouldIgnoreStatistics` via `ParquetRewriter.convertStatistics` | Pages per rewritten chunk | Rewrite |
| 4 | `CorruptStatistics.shouldIgnoreStatistics` via `EncryptedColumnChunkMetaData.decryptIfNeeded` | 1 per encrypted column | Lazy decrypt |
| 5 | `CorruptDeltaByteArrays.requiresSequentialReads(String, Encoding)` in `ParquetRecordReader` | 1 per reader init | Reader init |
| 6 | `ColumnReadStoreImpl` constructor via `MessageColumnIO.getRecordReader` | R (once per row group read) | Row group materialization |
| 7 | `ColumnReadStoreImpl` constructor via `ParquetRewriter.nullifyColumn` | 1 per nullified column | Rewrite with nullification |

`ColumnReadStoreImpl` already parses `createdBy` into a `ParsedVersion` and stores it as a field — but does so R times (once per row group) because nobody upstream caches the parsed result.

## Proposed Change

Add a cached `ParsedVersion` field to `FileMetaData`:

```java
// In FileMetaData:
private final transient ParsedVersion writerVersion;

public FileMetaData(...) {
...
this.createdBy = createdBy;
this.writerVersion = parseVersion(createdBy);
...
}

public ParsedVersion getWriterVersion() {
return writerVersion;
}

private static ParsedVersion parseVersion(String createdBy) {
if (Strings.isNullOrEmpty(createdBy)) return null;
try {
return VersionParser.parse(createdBy);
} catch (RuntimeException | VersionParseException e) {
return null;
}
}
```

Then add `ParsedVersion`-accepting overloads following the existing pattern established by `CorruptDeltaByteArrays.requiresSequentialReads(ParsedVersion, Encoding)`:

- `CorruptStatistics.shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName)`
- `ColumnReadStoreImpl` constructor accepting `ParsedVersion` directly

## Why This Approach

- **Purely additive**: new `transient` field + getter, no breaking changes
- **Doesn't break serialization**: field is `transient`
- **Follows existing precedent**: `CorruptDeltaByteArrays` already has `ParsedVersion`-based overloads used from `ColumnReaderBase`
- **Enables incremental adoption**: call sites can migrate to the cached version one at a time
- **Unblocks #3607**: PR #3607 can rebase onto this foundation cleanly using `ParsedVersion`-based APIs instead of threading a PARQUET-251-specific boolean

## Scope

This issue covers:
1. Adding the `ParsedVersion` field and getter to `FileMetaData`
2. Adding `shouldIgnoreStatistics(ParsedVersion, PrimitiveTypeName)` overload to `CorruptStatistics`
3. Updating `ColumnReadStoreImpl` to accept `ParsedVersion` directly
4. Migrating existing call sites to use the cached version where `FileMetaData` is accessible

## Context

Discussion: https://github.com/apache/parquet-java/pull/3607#issuecomment-5006760055
Related: #3601

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Beginne mit dem Lesen von FileMetaData, CorruptStatistics und ColumnReadStoreImpl und verfolge anschließend die aufgeführten Aufrufstellen von VersionParser.parse(createdBy) sowie die bestehende ParsedVersion-Überladung in CorruptDeltaByteArrays. Die Arbeit ist abgeschlossen, wenn FileMetaData die zwischengespeicherte geparste Version bereitstellt, die angeforderten Überladungen vorhanden sind und zugängliche Aufrufstellen den zwischengespeicherten Wert verwenden, ohne das Serialisierungsverhalten zu ändern.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
java
Bereich
data
Issue-Typ
Refactoring
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Ruhig
Klarheit
Klar beschrieben
Anfängerfreundlichkeit
55/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.