apache / apache/arrow-java

Out-of-bounds read for corrupt view offsets in BaseVariableWidthViewVector

未关闭
#1,217 0 条评论 0 个 reaction 已指派 0 人 已被 @lidavidm 认领 在 GitHub 查看
主要语言
Java
星标
94
派生
152
平均合并
3 天 16 小时
30 天内合并 PR
11

描述

### Describe the bug

`ViewVarCharVector`/`ViewVarBinaryVector` store values longer than `INLINE_SIZE` (12 bytes) out of line, encoding a data-buffer index and an offset inline in the view buffer. When a vector is loaded from an IPC stream these fields come straight from the input.

`BaseVariableWidthViewVector` dereferences them verbatim in `getData`, `getDataPointer`, `hashCode`, `copyFromNotNull` and `splitAndTransferViewBufferAndDataBuffer`, e.g.

```java
dataBuffers.get(bufferIndex).getBytes(dataOffset, result, 0, dataLength);
```

Nothing checks that `bufferIndex` is in range or that `dataOffset + dataLength` fits inside the referenced data buffer. A crafted view whose offset/length points past the data buffer produces an out-of-bounds read: with the default bounds checking it throws `IndexOutOfBoundsException`, but with `arrow.enable_unsafe_memory_access=true` (commonly set in production) it reads arbitrary native heap into the returned value.

### Component(s)

Java

贡献指南

打开贡献指南

调研方向

从 BaseVariableWidthViewVector 开始,检查 getData、getDataPointer、hashCode、copyFromNotNull 和 splitAndTransferViewBufferAndDataBuffer 中对 out-of-line view 的处理。对比关联的 pull request,然后验证损坏的 buffer 索引和 offset/length 范围不再允许越界读取,包括启用 unsafe memory access 的情况。

由索引模型根据 Issue 内容生成。

评估

技术栈
java
领域
security
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
描述清楚
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。