apache / apache/arrow

[C++] ReadRangeCache should not retain data after read

Open
#32,846 11 comments 0 reactions 0 assignees View on GitHub
Component: C++ good-second-issue Type: enhancement
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 18h
Merged PRs (30d)
91

Description

I've added a unit test of the issue here: https://github.com/westonpace/arrow/tree/experiment/read-range-cache-retention

We use the ReadRangeCache for pre-buffering IPC and parquet files. Sometimes those files are quite large (gigabytes). The usage is roughly:

for X in num_row_groups:
CacheAllThePiecesWeNeedForRowGroupX
WaitForPiecesToArriveForRowGroupX
ReadThePiecesWeNeedForRowGroupX

However, once we've read in row group X and passed it on to Acero, etc. we do not release the data for row group X. The read range cache's entries vector still holds a pointer to the buffer. The data is not released until the file reader itself is destroyed which only happens when we have finished processing an entire file.

This leads to excessive memory usage when pre-buffering is enabled.

This could potentially be a little difficult to implement because a single read range's cache entry could be shared by multiple ranges so we will need some kind of reference counting to know when we have fully finished with an entry and can release it.

**Reporter**: [Weston Pace](https://issues.apache.org/jira/browse/ARROW-17599) / @westonpace
**Assignee**: [Percy Camilo Triveño Aucahuasi](https://issues.apache.org/jira/browse/ARROW-17599) / @aucahuasi
**Watchers**: [Rok Mihevc](https://issues.apache.org/jira/browse/ARROW-17599) / @rok
#### Related issues:
- [[C++] Implement a read range process without caching](https://github.com/apache/arrow/issues/33311) (is related to)
- [Lower memory usage with filters](https://github.com/apache/arrow/issues/32838) (is related to)
#### PRs and other links:
- [GitHub Pull Request #14226](https://github.com/apache/arrow/pull/14226)

**Note**: *This issue was originally created as [ARROW-17599](https://issues.apache.org/jira/browse/ARROW-17599). Please see the [migration documentation](https://github.com/apache/arrow/issues/14542) for further details.*

Contributor guide

Open the contributing guide

Research direction

Start with the ReadRangeCache implementation and the unit test on the linked experiment branch. Trace how cache entries are retained across IPC or Parquet row-group reads, then verify that data is released after its final read while shared ranges remain valid; use PR #14226 and the related issues for existing work and context.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
data, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.