pydata / pydata/xarray

Datatype for a 'shape specification' of a Dataset / DataArray

Open
#6,680 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement topic-typing
Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

Is your feature request related to a problem?

Often with xarray I find myself having to create a template Dataset or DataArray with dummy data in it just to specify the dimensions/sizes/coordinates/variable names that are required in some situation.

Describe the solution you'd like

It would be very useful to have a datatype that represents a shape specification (dimensions, sizes and coordinates) independently of the data so that we can do things like:

  • Implement xarray equivalents of functions like np.ones, np.zeros, np.random.normal(size=...) that are given a shape specification which the return value should conform to. (I have some more sophisticated / less trivial examples of this too, functions which currently need to be given templates for the return value but only depend on the shape of the template)
  • Test if two DataArrays / Datasets have the same shape
  • Memoize or cache things based on shape (this implies the shape spec would need to be hashable)
  • Make it easier to use xarray with libraries like tree / PyTree that can be used to flatten and unflatten a Dataset into its underlying arrays together with some specification of the shape of the data structure that can be used to unflatten it back again. (Right now I have to implement my own shape specification objects to do this)
  • Manipulate shape specifications e.g. by adding or removing dimensions from them without having to manipulate dummy template data in slightly arbitrary ways (e.g. template.isel(dim_to_be_dropped=0, drop=True)) in order to do this.
Describe alternatives you've considered

I realise that using lazy dask arrays largely removes the performance overhead of manipulating fake data, but (A) it still feels kinda ugly and adds boilerplate to construct the fake data, and (B) not everyone wants to depend on dask.

Additional context

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing how template Dataset and DataArray objects currently express dimensions, sizes, coordinates, and variable names, including the isel(..., drop=True) pattern mentioned in the issue. Compare the requirements from np.ones, np.zeros, and np.random.normal(size=...), then define what a standalone specification must represent, how it can be manipulated and hashed, and how completion would support shape comparison and reconstruction.

Written by the indexing model from the issue text.

Assessment

Tech stack
numpy, python
Domain
data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.