apache / apache/parquet-java

record count for row group size check configurable

Abierto
#2,740 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Component: Java Component: Parquet Priority: Major Type: enhancement
Lenguaje dominante
Java
Estrellas
3.1k
Forks
1.6k
Merge medio
3 d 12 h
PR fusionados (30 d)
33

Descripción

 org.apache.parquet.hadoop.InternalParquetRecordWriter#checkBlockSizeReached
```java

 private void checkBlockSizeReached() throws IOException {
    if (recordCount >= recordCountForNextMemCheck) { // checking the memory size is relatively expensive, so let's not do it for every record.
      long memSize = columnStore.getBufferedSize();
      long recordSize = memSize / recordCount;
      // flush the row group if it is within ~2 records of the limit
      // it is much better to be slightly under size than to be over at all
      if (memSize > (nextRowGroupSize - 2 * recordSize)) {
        LOG.info("mem size {} > {}: flushing {} records to disk.", memSize, nextRowGroupSize, recordCount);
        flushRowGroupToStore();
        initStore();
        recordCountForNextMemCheck = min(max(MINIMUM_RECORD_COUNT_FOR_CHECK, recordCount / 2), MAXIMUM_RECORD_COUNT_FOR_CHECK);
        this.lastRowGroupEndPos = parquetFileWriter.getPos();
      } else {
        recordCountForNextMemCheck = min(
            max(MINIMUM_RECORD_COUNT_FOR_CHECK, (recordCount + (long)(nextRowGroupSize / ((float)recordSize))) / 2), // will check halfway
            recordCount + MAXIMUM_RECORD_COUNT_FOR_CHECK // will not look more than max records ahead
            );
        LOG.debug("Checked mem at {} will check again at: {}", recordCount, recordCountForNextMemCheck);
      }
    }
  }
```
in this code,if the block size is small ,for example 8M,and the first 100 lines record size is small and  after 100 lines the record size is big,it will cause big row group,in our real scene,it will more than 64M. So i think the size for block check can configurable.

**Reporter**: [xjlem](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=xjlem)

**Note**: *This issue was originally created as [PARQUET-2242](https://issues.apache.org/jira/browse/PARQUET-2242). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Comienza con org.apache.parquet.hadoop.InternalParquetRecordWriter#checkBlockSizeReached y rastrea cómo se inicializan y utilizan nextRowGroupSize y recordCountForNextMemCheck. Identifica la ruta de configuración existente y las pruebas relevantes del writer; después, define la finalización como permitir configurar el dimensionamiento de la comprobación de memoria y evitar el comportamiento reportado de los row groups sobredimensionados.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
java
Área
data-engineering
Tipo de issue
Nueva funcionalidad
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
42/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.