repr(sdata) reports the parquet row count for any points element that shares backing files, even after lazy filtering
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 394
- Forks
- 95
- Avg merge
- 4d 3h
- Merged PRs (30d)
- 7
Description
[!NOTE]
This whole message is AI-generated. The issue was automatically discovered and reported by an AI agent (Claude) during an autonomous bug hunt on thespatialdatacode base. It has not been verified or triaged by a human yet; theneeds: triagelabel is set so that a maintainer can confirm it. The reproduction script below was executed by the agent in an isolated environment (see Environment) and its output is pasted verbatim.
Summary
sdata.points["filtered"] = PointsModel.parse(sdata["pts"][sdata["pts"].x > 90]) has 9 rows but the repr shows DataFrame with shape: (100, 2).
Severity (agent's assessment): low/medium (misleading output)
Where: src/spatialdata/_core/spatialdata.py::_gen_repr (sums pq.read_metadata(f).num_rows over the Dask backing files)
Expected behaviour
The frame's own length (or <Delayed>).
Reproduction
Save as repro.py and run uv run repro.py (the PEP 723 header pins spatialdata to the commit the bug was found on; replace the URL fragment with @main to test the current main branch).
# /// script
# requires-python = ">=3.12"
# dependencies = [
# "spatialdata @ git+https://github.com/scverse/spatialdata.git@ccf1ea048d054b6624214bf618008a9f9ae223e0",
# ]
# ///
"""repr(sdata) reports the parquet row count for a lazily filtered points element that shares backing files."""
import os
import shutil
import tempfile
import warnings
import numpy as np
import pandas as pd
from spatialdata import SpatialData, read_zarr
from spatialdata.models import PointsModel
warnings.simplefilter("ignore")
tmp = tempfile.mkdtemp()
SpatialData(points={"pts": PointsModel.parse(pd.DataFrame({"x": np.arange(100.0), "y": np.arange(100.0)}))}).write(os.path.join(tmp, "store.zarr"))
sdata = read_zarr(os.path.join(tmp, "store.zarr"))
sdata.points["filtered"] = PointsModel.parse(sdata["pts"][sdata["pts"].x > 90])
n = len(sdata["filtered"].compute())
line = [l.strip() for l in repr(sdata).splitlines() if "'filtered'" in l][0]
print("actual rows of 'filtered':", n)
print("repr line :", line)
shutil.rmtree(tmp)
bug = f"({n}," not in line
print("VERDICT:", "BUG REPRODUCED" if bug else "NOT REPRODUCED")
Observed output
actual rows of 'filtered': 9
repr line : ├── 'filtered': DataFrame with shape: (100, 2) (2D points)
VERDICT: BUG REPRODUCED
Possible fix direction (unverified)
Only use the parquet metadata shortcut when the expression is a bare read_parquet; otherwise fall back to <Delayed> or len(v).
Environment
uv run repro.py with the PEP 723 metadata in the script (fresh, isolated environment; spatialdata built from main @ ccf1ea0 (2026-08-28); Python 3.13, latest releases of the dependencies at run time: pandas 3.0, anndata 0.13, zarr 3.3, dask 2026.8, numpy 2.5, geopandas 1.1, shapely 2.1). macOS (arm64). Also reproduced in a second environment with pandas 2.3.3 / anndata 0.12.11 / numpy 2.4.4 / zarr 3.2.1.
Automatically generated; discovered by an AI agent (Claude) and not yet reviewed by a human.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with src/spatialdata/_core/spatialdata.py::_gen_repr and run the provided repro.py with uv run to confirm the filtered points element reports 9 rows while sharing backing files. Trace how the representation derives its row count for the Dask-backed expression. Done means repr reports the filtered frame’s length or instead of the parquet backing-file count.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Active
- Clarity
- Clearly specified
- Newbie friendliness
- 76/100