Allow for custom compression codecs
- Lenguaje dominante
- Java
- Estrellas
- 3.1k
- Forks
- 1.6k
- Merge medio
- 3 d 12 h
- PR fusionados (30 d)
- 33
Descripción
I understand that the list of accepted compression codecs is explicity limited to uncompressed, snappy, gzip, and lzo. (See parquet.hadoop.metadata.CompressionCodecName.java) Is there a reason for this? Or is there an easy workaround? On the surface it seems like an unnecessary restriction.
I ask because I have written a custom codec to implement encryption and I'm unable to use it with Parquet, which is a real shame because it is the main storage format I was hoping to use.
Other thoughts on how to implement encryption in Parquet with this limitation?
**Reporter**: [Steven Anton](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=santon)
**Note**: *This issue was originally created as [PARQUET-678](https://issues.apache.org/jira/browse/PARQUET-678). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
Start with parquet.hadoop.metadata.CompressionCodecName.java, the file named in the issue, and review the migration documentation linked in the issue note. Clarify whether the goal is custom codec support, an encryption workaround, or both; done should be a decided and documented approach that permits the requested use case.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- java
- Área
- data-engineering
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Estancado
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 25/100