apache / apache/arrow-java

[Java][C] sliced RecordBatch offset info is lost when imported from c-data

オープン
#88 コメント 6 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Java
スター
94
フォーク
152
平均マージ
3日 16時間
マージ済み PR(30日)
11

説明

### Describe the bug, including details regarding any error messages, version, and platform.

Reproduced on latest arrow release (16.0)

When importing a sliced RecordBatch from c to java

On c side:
```
auto sliced_record_batch = original_record_batch->Slice(/*offset=*/8, /*length=*/2);
arrow::ExportRecordBatch(sliced_record_batch, arrow_array_ptr);
```

On java side:
```
ArrowArray arrowArray = ArrowArray.allocateNew(allocator);
Data.importIntoVectorSchemaRoot(allocator, arrowArray, vectorSchemaRoot, null);
```

The imported vectorSchemaRoot maintains the correct length(which is 2), but the offset info (which is 8) is not respected, hence the content of the imported vectorSchemaRoot points to the first 2 rows of the original_record_batch, while the desired content is sliced_record_batch.

I'm not familiar with arrow code, but it seems that the offset info is actually present in org.apache.arrow.c.ArrowArray.Snapshot, but org.apache.arrow.c.ArrayImporter ignores the offset in org.apache.arrow.c.ArrayImporter.doImport(ArrowArray.Snapshot)

### Component(s)

Java

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

まず org.apache.arrow.c.ArrayImporter.doImport(ArrowArray.Snapshot) と org.apache.arrow.c.ArrowArray.Snapshot の offset 情報を読みます。C 側でスライスした RecordBatch を使って Java のインポートを再現し、インポートされた vectorSchemaRoot に元の batch の先頭行ではなく、スライスの offset からの行が含まれていることを確認します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
c, java
領域
data-engineering
issue の種類
バグ
難易度
3/5
見積もり時間
1〜2日
活発さ
停滞
明瞭さ
明確に書かれている
初心者へのやさしさ
45/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。