Add batch/columnar write API for Arrow VectorSchemaRoot (Java parity with C++/Python)
- Langage dominant
- Java
- Étoiles
- 3.1k
- Forks
- 1.6k
- Merge moyen
- 3 j 12 h
- PR mergées (30 j)
- 33
Description
### Problem
`parquet-java` has no public API to write Arrow `VectorSchemaRoot` to Parquet files. Callers must construct row objects and feed them through `ParquetWriter.write(T)` one at a time. Arrow C++ and PyArrow support this natively via `parquet::WriteTable()` / `pq.write_table()`.
Related prior discussion:
- #2264 (open since 2019)
- #3353
### Motivation
Several downstream Java projects work with Arrow-columnar data internally and must materialize row objects solely to satisfy `ParquetWriter`'s input API:
- Apache Iceberg (`FileAppender`) — [iceberg#17748](https://github.com/apache/iceberg/issues/17748)
- Apache Fluss (Arrow-native streaming storage) — [fluss#4047](https://github.com/apache/fluss/issues/4047)
- Apache Paimon (worked around this by building their own `paimon-arrow` writer that bypasses `parquet-java`'s row API entirely)
### Key Insight
For PLAIN-encoded, fixed-width columns, Arrow's in-memory format (contiguous little-endian values) is identical to Parquet's PLAIN page encoding. A zero-copy path is possible by wrapping the Arrow data buffer directly as a `BytesInput` and passing it to `PageWriter.writePage()`.
For nullable columns, Arrow's validity bitmap can be scanned for contiguous non-null runs, with each run bulk-copied and definition levels encoded as RLE runs of the same value — O(null_transitions) instead of O(N).
### Proposed Design
A new `ArrowParquetWriter` in the `parquet-arrow` module that writes pages directly to `PageWriter` rather than going through `RecordConsumer`/`ColumnWriter`:
```
ArrowParquetWriter.writeBatch(VectorSchemaRoot)
→ per column: selects optimal write strategy
→ writes pages to PageWriter via getPageWriter(ColumnDescriptor)
→ manages row groups via ParquetFileWriter
```
Per-column strategy selection (best to worst):
1. **Zero-copy** (non-null + fixed-width + PLAIN): wrap Arrow buffer as `BytesInput` directly. RL/DL as single-value RLE runs. Stats from sequential buffer scan.
2. **Bulk-copy with nulls** (nullable + fixed-width + PLAIN): scan validity bitmap for non-null runs, bulk-copy each, emit DL as RLE runs per null-transition.
3. **Variable-width rewrite** (PLAIN + string/binary): single-pass transformation of Arrow offset+data buffers to Parquet's length-prefixed format.
4. **Dictionary**: map Arrow's `DictionaryEncodedVector` to Parquet dictionary pages, or build dictionary from plain vector.
5. **Fallback**: per-value through `ValuesWriter` (same cost as today, for unsupported encoding/type combos).
All verified public APIs exist for this:
- `PageWriteStore.getPageWriter(ColumnDescriptor)` — get column page writer directly
- `PageWriter.writePage(BytesInput, valueCount, rowCount, stats, encodings)` — write pre-encoded pages
- `BytesInput.from(ByteBuffer)` — zero-copy buffer wrapping
- `RunLengthBitPackingHybridEncoder` — encode RL/DL levels
- `ColumnChunkPageWriteStore.flushToFileWriter()` — flush pages to file
### Why Not Extend ParquetWriter?
`ParquetWriter.write(T)` calls `InternalParquetRecordWriter.write()` which increments `recordCount` by 1 per call and checks row-group boundaries based on that count. This is incompatible with batch semantics. The writer must manage `ParquetFileWriter` and row groups directly.
### Scope
- Parquet file format unchanged
- Existing `ParquetWriter` API unchanged
- `parquet-arrow` gains `parquet-hadoop` as a compile dependency (for `ParquetFileWriter`, `ColumnChunkPageWriteStore`)
- No Hadoop runtime dependency for uncompressed output (compression requires caller-supplied `CompressionCodecFactory`)
- Flat schemas (primitive columns) first; nested types deferred
### Implementation Phases
1. Core infrastructure + zero-copy for non-null fixed-width PLAIN columns
2. Nullable column support (validity bitmap scanning + run-based bulk copy)
3. Variable-width types (string/binary offset rewriting)
4. Dictionary encoding support
5. Bloom filter, nested types, advanced features
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Piste de recherche
Commencez par lire le module `parquet-arrow` et les API nommées `PageWriteStore`, `PageWriter`, `BytesInput` et `ColumnChunkPageWriteStore`. Le travail proposé couvre cinq phases d’implémentation, en commençant par les écritures zero-copy pour les colonnes PLAIN de largeur fixe non nullables ; l’achèvement doit être évalué par rapport à la phase choisie, la portée complète incluant la prise en charge des colonnes nullable, de largeur variable et des dictionnaires.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- java
- Domaine
- data-engineering
- Type d'issue
- Fonctionnalité
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Activité
- Active
- Clarté
- Plutôt claire
- Accessibilité débutants
- 30/100