apache / apache/arrow-java

[Java] Unexpected RecordBatch length when saving empty table to file with compression

未關閉
#194 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
Type: bug
主要語言
Java
星號
94
分支
152
平均合併
3 天 16 小時
30 天內合併 PR
11

描述

### Describe the bug, including details regarding any error messages, version, and platform.

This might be more of a usage question since I couldn't find anything in the format docs on how to set the length field with compression.

The issue is that if I try to read an empty table with the [Julia extension](https://github.com/apache/arrow-julia) it just hangs. The reason for this seems to be that it [only checks](https://github.com/apache/arrow-julia/blob/e893c327f177f5a4d5efeab831df0fe93ab4ec5b/src/table.jl#L518-L529) the length field in the RecordBatch when deciding whether to attempt to decode and not the length read from the first 8 bytes of the data.

The file created by the code below is readable by both pyarrow and the java implementation, so chances are that the Julia implementation is doing it wrong (I will open an issue there as well). Is there some reference to how one shall interpret the length field in RecordBatch when using compression?

Code to create an empty table in case I'm doing something wrong

```java
public static void main(String[] args) {
try (BufferAllocator allocator = new RootAllocator()) {
Field name = new Field("name", FieldType.nullable(new ArrowType.Utf8()), null);
Field age = new Field("age", FieldType.nullable(new ArrowType.Int(32, true)), null);
Schema schemaPerson = new Schema(asList(name, age));
try(
VectorSchemaRoot vectorSchemaRoot = VectorSchemaRoot.create(schemaPerson, allocator)
){
vectorSchemaRoot.allocateNew(); // Needed?
vectorSchemaRoot.setRowCount(0); // Needed?
File file = new File("randon_access_to_file.arrow");
try (
FileOutputStream fileOutputStream = new FileOutputStream(file);
ArrowFileWriter writer = new ArrowFileWriter(vectorSchemaRoot, null, fileOutputStream.getChannel(),
null, IpcOption.DEFAULT,
CommonsCompressionFactory.INSTANCE, CompressionUtil.CodecType.ZSTD)
) {
writer.start();
writer.writeBatch();
writer.end();
System.out.println("Record batches written: " + writer.getRecordBlocks().size() + ". Number of rows written: " + vectorSchemaRoot.getRowCount());
} catch (IOException e) {
e.printStackTrace();
}
}
}
}
```

When I tried saving a compressed empty table using pyarrow I got 0 as the length field and the Julia implementation could read the table without hanging.

Disclaimer: I don't have a working python installation so I did this though PythonCall. Hopefully I managed to remove all the Julia-isms so that it runs in python:
```python
schema = pa.schema([pa.field('nums', pa.int32())])

with pa.OSFile('bigfile.arrow', 'wb') as sink:
with pa.ipc.new_file(sink, schema, options=pa.ipc.IpcWriteOptions(compression='zstd'))) as writer:
batch = pa.record_batch([pa.array([], type=pa.int32())], schema)
writer.write(batch)
```

### Component(s)

Java

貢獻指南

開啟貢獻指南

研究方向

先檢查 Java 的 ArrowFileWriter 以及範例使用的 IPC 壓縮處理,接著將其空的壓縮 RecordBatch 與 pyarrow 的輸出和 Julia reader 的 table.jl 邏輯進行比較。確認預期的長度欄位解讀方式,並新增涵蓋空壓縮 batch 的回歸測試;當產生的檔案能夠互通,且 reader 不會卡住時,即表示完成。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
java
領域
data-engineering
Issue 類型
缺陷
難度
4/5
預估耗時
3-5 天
活躍度
停滯
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。