pydata / pydata/xarray

wrong chunk sizes in html repr with nonuniform chunks

Open
#4,376 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

topic-html-repr
Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

What happened:

The HTML repr is using the first element in a chunks tuple;

What you expected to happen:

it should be using whatever dask does in this case

Minimal Complete Verifiable Example:


import xarray as xr
import dask


test = xr.DataArray(
    dask.array.zeros(
        (12, 901, 1001),
        chunks=(
            (1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1),
            (1, 899, 1),
            (1, 199, 1, 199, 1, 199, 1, 199, 1, 199, 1),
        ),
    )
)
test.to_dataset(name="a")

image

EDIT: The text repr has the same issue

<xarray.Dataset>
Dimensions:  (dim_0: 12, dim_1: 901, dim_2: 1001)
Dimensions without coordinates: dim_0, dim_1, dim_2
Data variables:
    a        (dim_0, dim_1, dim_2) float64 dask.array<chunksize=(1, 1, 1), meta=np.ndarray>

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the issue's Python example and inspect the DataArray and Dataset HTML and text representations, focusing on how nonuniform chunk metadata is displayed. Compare the reported chunk sizes with the original chunks tuple; done means both representations show the correct chunk sizes for every dimension.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.