apache / apache/parquet-java

Add batch/columnar write API for Arrow VectorSchemaRoot (Java parity with C++/Python)

Aperta
#3,733 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Java
Stelle
3.1k
Fork
1.6k
Merge medio
3g 12h
PR unite (30g)
33

Descrizione

### Problem

`parquet-java` has no public API to write Arrow `VectorSchemaRoot` to Parquet files. Callers must construct row objects and feed them through `ParquetWriter.write(T)` one at a time. Arrow C++ and PyArrow support this natively via `parquet::WriteTable()` / `pq.write_table()`.

Related prior discussion:
- #2264 (open since 2019)
- #3353

### Motivation

Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy `ParquetWriter`'s input API:

- Apache Iceberg (`FileAppender`) — [iceberg#17748](https://github.com/apache/iceberg/issues/17748)
- Apache Fluss (Arrow-native streaming storage) — [fluss#4047](https://github.com/apache/fluss/issues/4047)
- Apache Paimon (worked around this by building their own `paimon-arrow` writer that bypasses `parquet-java`'s row API entirely)

### Key Insight

For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a `BytesInput` and passing it to `PageWriter.writePage()`.

For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).

### Proposed Design

A new `ArrowParquetWriter` in the `parquet-arrow` module that writes pages directly to `PageWriter` rather than going through `RecordConsumer`/`ColumnWriter`:

```
ArrowParquetWriter.writeBatch(VectorSchemaRoot)
→ per column: selects optimal write strategy
→ writes pages to PageWriter via getPageWriter(ColumnDescriptor)
→ manages row groups via ParquetFileWriter
```

Per-column strategy selection (best to worst):

1. **Zero-copy** (non-null + fixed-width + PLAIN): wrap Arrow buffer as `BytesInput` directly. RL/DL as single-value RLE runs. Stats from sequential buffer scan.
2. **Bulk-copy with nulls** (nullable + fixed-width + PLAIN): scan validity bitmap for non-null runs, bulk-copy each, emit DL as RLE runs per null-transition.
3. **Variable-width rewrite** (PLAIN + string/binary): single-pass transformation of Arrow offset+data buffers to Parquet's length-prefixed format.
4. **Dictionary**: map Arrow's `DictionaryEncodedVector` to Parquet dictionary pages, or build dictionary from plain vector.
5. **Fallback**: per-value through `ValuesWriter` (same cost as today, for unsupported encoding/type combos).

All verified public APIs exist for this:
- `PageWriteStore.getPageWriter(ColumnDescriptor)` — get column page writer directly
- `PageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings)` — write pre-encoded pages
- `BytesInput.from(ByteBuffer)` — zero-copy buffer wrapping
- `RunLengthBitPackingHybridEncoder` — encode RL/DL levels
- `ColumnChunkPageWriteStore.flushToFileWriter()` — flush pages to file

### Why Not Extend ParquetWriter?

`ParquetWriter.write(T)` calls `InternalParquetRecordWriter.write()` which increments `recordCount` by 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manage `ParquetFileWriter` and row groups directly.

### Scope

- Parquet file format unchanged
- Existing `ParquetWriter` API unchanged
- `parquet-arrow` gains `parquet-hadoop` as a compile dependency (for `ParquetFileWriter`, `ColumnChunkPageWriteStore`)
- No Hadoop runtime dependency for uncompressed output (compression requires caller-supplied `CompressionCodecFactory`)
- Flat schemas (primitive columns) first; nested types deferred

### Implementation Phases

1. Core infrastructure + zero-copy for non-null fixed-width PLAIN columns
2. Nullable column support (validity bitmap scanning + run-based bulk copy)
3. Variable-width types (string/binary offset rewriting)
4. Dictionary encoding support
5. Bloom filter, nested types, advanced features

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

Inizia leggendo il modulo `parquet-arrow` e le API indicate `PageWriteStore`, `PageWriter`, `BytesInput` e `ColumnChunkPageWriteStore`. Il lavoro proposto comprende cinque fasi di implementazione, iniziando dalle scritture zero-copy per colonne PLAIN a larghezza fissa non nullable; il completamento dovrebbe essere valutato in base alla fase scelta, mentre l'ambito completo include il supporto per colonne nullable, a larghezza variabile e per i dizionari.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
java
Ambito
data-engineering
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Attiva
Chiarezza
Abbastanza chiara
Idoneità per principianti
30/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.