apache / apache/arrow

[C++][Python] Default value for CompressedInputStream kChunkSize might be too small

Open
#41,604 14 comments 0 reactions 1 assignee Claimed by @pitrou View on GitHub
Component: C++ Component: Python Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

### Describe the bug, including details regarding any error messages, version, and platform.

I'm using pyarrow.csv.open_csv to stream read a 15GB gz csv file over S3. The speed is unusable slow.

```python

file_src = pyarrow_s3fs.open_input_stream(path_src)

read_options = pyarrow.csv.ReadOptions(block_size=5_000_000, encoding="latin1")
csv = pyarrow.csv.open_csv(
file_src,
read_options=read_options,
parse_options=parse_options,
convert_options=convert_options,
)
for batch in self._csv:
...
```

Here, even when I set the block_size=5_000_000, the reader are issuing 65K ranged read over S3.

This is bad for two reason:
1. the S3 has cost for each get request ($0.0004/1k req)
2. compute the authentication header etc has computation cost

After digging into the code, I find this `kChunkSize` is hard coded in CompressedInputStream https://github.com/apache/arrow/blob/f6127a6d18af12ce18a0b8b1eac02346721cc399/cpp/src/arrow/io/compressed.cc#L432

Currently, my workaround is to use buffered stream
`file_src = pyarrow_s3fs.open_input_stream(path_src, buffer_size=10_000_000)`

But it is not obvious from the doc. Could we set this value higher? Or at least add some doc to clarify the usage.

### Component(s)

C++, Python

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.