apache / apache/arrow-java

DictionaryEncoder.decode accepts out-of-range dictionary indices

未關閉 適合新手
#1,261 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Java
星號
94
分支
152
平均合併
3 天 16 小時
30 天內合併 PR
11

描述

`DictionaryEncoder.retrieveIndexVector` guards each index from the index vector with `indexAsInt > dictionaryCount` before `transfer.copyValueSafe(indexAsInt, i)`. Valid indices are `0..dictionaryCount-1`, so the check is off by one: an index equal to `dictionaryCount` is accepted and reads one slot past the dictionary vector, and a negative index (a signed index type with the high bit set) is not rejected either and also reaches `copyValueSafe`. The index vector is decoded from an IPC/C-data payload, so a crafted dictionary-encoded batch yields an out-of-bounds read of the dictionary vector, exposing adjacent off-heap memory when bounds checking is disabled via `arrow.enable_unsafe_memory_access`.

The same helper backs `DictionaryEncoder.decode`, `ListSubfieldEncoder.decodeListSubField` and `StructSubfieldEncoder.decode`.

The bound should be `indexAsInt < 0 || indexAsInt >= dictionaryCount`.

貢獻指南

開啟貢獻指南

研究方向

從 DictionaryEncoder.retrieveIndexVector 開始,追蹤其從 DictionaryEncoder.decode、ListSubfieldEncoder.decodeListSubField 和 StructSubfieldEncoder.decode 中的使用。驗證負索引以及等於 dictionaryCount 的索引是否會在 transfer.copyValueSafe 之前被拒絕;完成標準是,構造的字典索引無法以無效位置到達字典副本。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
java
領域
security
Issue 類型
缺陷
難度
2/5
預估耗時
1-3 小時
活躍度
活躍
描述清晰度
描述清楚
新手友好度
74/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。