pydata / pydata/xarray

`_decode_cf_datetime_dtype` triggers 2 store reads per time variable — expensive on remote stores

Open
#11,303 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

Problem

_decode_cf_datetime_dtype (xarray/coding/times.py:339-371) reads the first and last element of every time-encoded variable during open_dataset / open_datatree to infer the decoded dtype. On local backends these reads are memory-mapped and free; on remote stores (zarr + S3, zarr + icechunk) each read is a round-trip through zarr.sync() → asyncio, even when chunks are already cached in the backend.

Cost scales as 2 × N_time_variables reads per open.

Downstream impact: this hits any workflow that opens many time-encoded groups from remote zarr — radar (CfRadial2: 1 time per sweep, 100+ sweeps per volume), CMIP model archives, satellite time-series, Pangeo-style DataTrees. Interactive notebooks, dashboards, and tile servers pay the cost on every open.

Impact — cProfile on main, 107-group DataTree (nexrad-arco/KLOT icechunk, 106 time variables)

total open_datatree time: 77.56s
Function cumtime share
_decode_cf_datetime_dtype 35.17s 45%
└─ first_n_items 25.16s 32%
└─ last_item 9.51s 12%
└─ decode_cf_datetime (actual decode) 0.09s 0.1%

The CPU cost of CF decoding is negligible. The 35 seconds are 100% I/O round-trip overhead.

Reproducer (public icechunk store, anonymous S3)

import icechunk, xarray as xr
from time import time

storage = icechunk.s3_storage(
    bucket="nexrad-arco", prefix="KLOT", region="us-east-1", anonymous=True,
)
repo = icechunk.Repository.open(storage)
session = repo.readonly_session("main")

start = time()
dtree = xr.open_datatree(session.store, engine="zarr", zarr_format=3,
                         consolidated=False, chunks={})
print(f"open_datatree: {time() - start:.1f}s")

Why the reads exist

The comment at times.py:346 (2018) explains: "Verify that at least the first and last date can be decoded successfully. Otherwise, tracebacks end up swallowed by Dataset.__repr__."

Empirically, no longer reproducible on current xarray — modern repr returns "..." for lazy arrays without triggering decode. But there is a second, undocumented purpose: with use_cftime=None and chunked data, decode_cf_datetime may return either datetime64[ns] or object (cftime) depending on whether values overflow the pandas ns range, and dask's map_blocks(func, array, dtype=dtype) casts output to the declared dtype — so a wrong declaration silently corrupts overflow values.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in xarray/coding/times.py:339-371, especially _decode_cf_datetime_dtype and its first_n_items and last_item calls. Reproduce the behavior with the 107-group icechunk KLOT DataTree example, then inspect how decode_cf_datetime chooses datetime64 versus object for chunked data. Done means reducing unnecessary remote reads while preserving correct overflow handling and dtype selection.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, numpy, pandas, python
Domain
data, performance
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
50/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.