[Java] Read a Parquet file into a Table
- Vorherrschende Sprache
- Java
- Sterne
- 94
- Forks
- 152
- Ø Merge
- 3 T. 16 Std.
- Gemergte PRs (30 T.)
- 11
Beschreibung
### Describe the usage question you have. Please include as many useful details as possible.
In Java, what is the canonical way of reading the whole content of a Parquet file into a Table?
I can read the file, as suggested, in a streaming fashion with VectorSchemaRoot, but then I miss how to collate all the batches into a single table.
From what I have understood, the options I have are:
1. build a big VectorSchemaRoot using VectorSchemaRootAppender, then invoke the Table constructor passing the vsr
2. construct FieldVectors explicitly, then read the parquet rows one by one (filling the FieldVectors as I go), then invoke the Table constructor passing a List
I see the C++ implementation has a convenient FromRecordBatches method.
### Component(s)
Java
Beitragsleitfaden
Rechercherichtung
Beginne damit, die Java-Verwendung von VectorSchemaRoot, VectorSchemaRootAppender und den Table-Konstruktoren zu lesen, und vergleiche sie anschließend mit der in der Issue erwähnten C++-Implementierung FromRecordBatches. Als abgeschlossen gilt die Arbeit, wenn das Projekt einen festgelegten, dokumentierten kanonischen Weg zum Laden einer vollständigen Parquet-Datei in eine Table oder eine gleichwertige Java-API mit Abdeckung hat.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- java
- Bereich
- data
- Issue-Typ
- Feature
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 30/100