Arrow vector in Java(Scala) allocate byteBuffer error while read the bytes from Python pyarrow
- Langage dominant
- Java
- Étoiles
- 94
- Forks
- 152
- Merge moyen
- 3 j 16 h
- PR mergées (30 j)
- 11
Description
I am using scala arrow 1.0.1 and pyarrow 1.0.1
Following error occurs when scala decode the byte that encoded from python.
tried to downgrade pyarrow to 0.17.0, 0.14.1, error still exists.
This is a loop function to decode the same data repeatedly. The error occurs sometimes, and I use a try catch block to include the code. And after an error occurs, following code could still work correctly, but after the first error occurs, following error would happen more frequently.
> Error stack trace java.nio.ByteBuffer.allocate(ByteBuffer.java:334)
> com.intel.analytics.zoo.shaded.arrow.vector.ipc.message.MessageSerializer.readMessage(MessageSerializer.java:692)
> com.intel.analytics.zoo.shaded.arrow.vector.ipc.message.MessageChannelReader.readNext(MessageChannelReader.java:57)
> com.intel.analytics.zoo.shaded.arrow.vector.ipc.ArrowStreamReader.readSchema(ArrowStreamReader.java:164)
> com.intel.analytics.zoo.shaded.arrow.vector.ipc.ArrowReader.initialize(ArrowReader.java:170)
> com.intel.analytics.zoo.shaded.arrow.vector.ipc.ArrowReader.ensureInitialized(ArrowReader.java:161)
> com.intel.analytics.zoo.shaded.arrow.vector.ipc.ArrowReader.getVectorSchemaRoot(ArrowReader.java:63)
How to fix it?
**Reporter**: [Litchy Soong](https://issues.apache.org/jira/browse/ARROW-10342)
**Note**: *This issue was originally created as [ARROW-10342](https://issues.apache.org/jira/browse/ARROW-10342). Please see the [migration documentation](https://github.com/apache/arrow/issues/14542) for further details.*
Guide de contribution
Ouvrir le guide de contribution
Piste de recherche
Commencez par reproduire la boucle de décodage répétée avec des octets produits par Python, puis suivez la pile depuis ArrowReader et ArrowStreamReader jusqu’à MessageChannelReader et MessageSerializer. L’issue ne nomme aucun fichier source ni aucun test ; le travail est considéré comme terminé lorsque les données peuvent être décodées de manière répétée sans l’échec intermittent d’allocation de ByteBuffer.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- java, python, scala
- Domaine
- data-engineering
- Type d'issue
- Bug
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Activité
- À l'abandon
- Clarté
- À clarifier
- Accessibilité débutants
- 25/100