apache / apache/iceberg-python

`DataScan` `count` method does not respect limit

Open
#2,121 10 comments 1 reaction 0 assignees View on GitHub
stale
Dominant language
Python
Stars
1.1k
Forks
581
Avg merge
1d 17h
Merged PRs (30d)
77

Description

### Apache Iceberg version

0.9.1 (latest release)

### Please describe the bug 🐞

When calling `count()` on a `DataScan`, limit is not respected. Seems trivial but if I set a limit of 5 I expect 5 or less rows back, at least with a scan-like implementation

The underlying `ArrowScan` does not get passed the limit param

https://github.com/apache/iceberg-python/blob/main/pyiceberg/table/__init__.py#L1940

This results in scans taking longer due to not respecting the limit.

The fix will involve more than just passing the limit to the `ArrowScan`

### Willingness to contribute

- [x] I can contribute a fix for this bug independently
- [ ] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [ ] I cannot contribute a fix for this bug at this time

Contributor guide

No contributing guide indexed for this repository

Research direction

Start in pyiceberg/table/__init__.py around line 1940 and trace DataScan.count through the underlying ArrowScan call. Check how the limit is propagated and identify the additional scan behavior the issue notes may need adjustment. Done means count() returns no more than the requested limit while avoiding unnecessary scan work.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering, databases
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.