apache / apache/arrow

[Python] File reading regression

Open
#29,772 2 comments 0 reactions 0 assignees View on GitHub
Component: Python Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

After ARROW-6626 was merged, we're seeing a slow down in file reading benchmarks:

One example: https://conbench.ursa.dev/benchmarks/b92e91fe8a4041148360d9433552277b/

**Reporter**: [Jonathan Keane](https://issues.apache.org/jira/browse/ARROW-14187) / @jonkeane
#### PRs and other links:
- [GitHub Pull Request #11284](https://github.com/apache/arrow/pull/11284)

**Note**: *This issue was originally created as [ARROW-14187](https://issues.apache.org/jira/browse/ARROW-14187). Please see the [migration documentation](https://github.com/apache/arrow/issues/14542) for further details.*

Contributor guide

Open the contributing guide

Research direction

Start with the linked Conbench benchmark and compare the file-reading results before and after ARROW-6626. Review GitHub pull request #11284 for the investigation and proposed resolution; done means identifying and resolving the regression so the affected benchmark returns to its prior performance.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, python
Domain
performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.