Slow performance of isel
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 4.2k
- Forks
- 1.4k
- Avg merge
- 2d 15h
- Merged PRs (30d)
- 14
Description
Hi,
I get a very slow performance of Dataset.isel or DataArray.isel in comparison with the native numpy approach. Do you know where this comes from?
ds = xr.Dataset(
{
"a": ("time", np.arange(55_000_000))
}, coords={
"time": np.arange(55_000_000)
}
)
time_filter = ds.time > 50_000
Select some values with DataArray.isel:
%timeit ds.a.isel(time=time_filter)
2.22 s ± 375 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
Use the native numpy approach:
%timeit ds.a.values[time_filter]
163 ms ± 12.1 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)
xarray: 0.10.4
pandas: 0.23.0
numpy: 1.14.2
scipy: 1.1.0
netCDF4: 1.4.0
h5netcdf: 0.5.1
h5py: 2.8.0
Nio: None
zarr: None
bottleneck: 1.2.1
cyordereddict: None
dask: 0.17.5
distributed: 1.21.8
matplotlib: 2.2.2
cartopy: 0.16.0
seaborn: 0.8.1
setuptools: 39.1.0
pip: 9.0.3
conda: None
pytest: 3.5.1
IPython: 6.4.0
sphinx: 1.7.4
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Reproduce the comparison between Dataset.isel or DataArray.isel and the native ds.a.values[time_filter] approach using the provided 55-million-element example. Trace the isel entry point and benchmark the relevant indexing path; done means the source of the slowdown is identified and a measurable performance improvement is demonstrated.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- numpy, python
- Domain
- data, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100