apache / apache/parquet-java

Parquet File not readable by Google big query (works with Spark)

Abierto
#2,550 8 comentarios 0 reacciones 0 asignados Ver en GitHub
Component: Avro Component: Parquet Priority: Blocker Type: bug
Lenguaje dominante
Java
Estrellas
3.1k
Forks
1.6k
Merge medio
3 d 12 h
PR fusionados (30 d)
33

Descripción

Hi
I'm trying to write Avro message to parquet on GCS. These parquet should be query by big query engine who support now parquet.

To do this I'm using Secor a kafka log persister tools from pinterest.

First I didn't notice any problem using Spark the same file can be read without any problem all is working perfect.

Now using Big query bring and error like this :
Error while reading table: , error message: Read less values than expected: Actual: 29333, Expected: 33827. Row group: 0, Column: , File:

After investigation using parquet-tools I figured out that in parquet there is metadata regarding number total of unique values for each columns eg from parquet-tools
page 0: DLE:BIT_PACKED RLE:BIT_PACKED [more]... CRC:[PAGE CORRUPT] VC:547

So the VC value indicate that the total number of unique value in the file is 547.

Now when make a spark SQL like SELECT DISTINCT COUNT(column) FROM ... I get 421 mean this number in the metadata is incorrect.

So what is not a problem for Spark to read is a blocking problem for Big data because it relies on these values and found it incorrect.

Is there any configuration of the writer that can prevent these errors in the metadata ? Or is it a normal behavior that should be a problem ?

Thanks

**Environment**: [secor|https://github.com/pinterest/secor]

GCP 

Big Query google cloud

Parquet writer 1.11

 

 
**Reporter**: [Richard Grossman](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=richiesgr)

**Note**: *This issue was originally created as [PARQUET-1946](https://issues.apache.org/jira/browse/PARQUET-1946). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

No se nombra ningún archivo del repositorio ni ninguna prueba. Empieza reproduciendo la salida de Secor con Parquet writer 1.11 y leyéndola en BigQuery; después, compara los metadatos indicados con los resultados de Spark y parquet-tools. Se considera terminado cuando se haya identificado la incompatibilidad y se haya confirmado con un caso de regresión específico o una limitación documentada.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
gcp, google-cloud, java
Área
data-engineering, databases
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Necesita aclaración
Aptitud para principiantes
30/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.