pydata / pydata/xarray

Explicit indexes: next steps

Open
#6,293 3 comments 12 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

topic-indexing
Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

#5692 is not merged yet now merged but and we can already start thinking about the next steps. I’m opening this issue to list and track the remaining tasks. @pydata/xarray, do not hesitate to add a comment below if you think about something that is missing here.

Continue the refactoring of the internals

Although in #5692 everything seems to work with the current pandas index wrappers for dimension coordinates, not all of Xarray's internals have been refactored yet to fully support (or at least be compatible with) custom indexes. Here is a list of Dataset / DataArray methods that still need to be checked / updated (this list may be incomplete):

  • as_numpy (#8001)
  • broadcast (#6430, #6481 )
  • drop_sel (#6605, #7699)
  • drop_isel
  • drop_dims
  • drop_duplicates (#8499)
  • transpose
  • interpolate_na
  • ffill
  • bfill
  • reduce
  • map
  • apply
  • quantile
  • rank
  • integrate
  • cumulative_integrate
  • filter_by_attrs
  • idxmin
  • idxmax
  • argmin
  • argmax
  • concat (partially refactored, may not fully work with multi-dimension indexes)
  • polyfit

I ended up following a common pattern in #5692 when adding explicit / flexible index support for various features (it is quite generic, though, the actual procedure may vary from one case to another and many steps may be skipped):

  • Check if it’s worth adding a new method to the Xarray Index base class. There may be several motivations:
    • Avoid handling Pandas index objects inside Dataset or DataArray methods (even if we don’t plan to fully support custom indexes for everything, it is preferable to put this logic behind the PandasIndex or PandasMultiIndex wrapper classes for clarity and also if eventually we want to make Xarray less dependent on Pandas)
    • We want a specific implementation rather than relying on the Variable’s corresponding method for speed-up or for other reasons, e.g.,
      • IndexVariable.concat exists to avoid unnecessary Pandas/Numpy conversions ; in #5692 PandasIndex.concat has the same logic and will fully replace the former if/once we get rid of IndexVariable
      • PandasIndex.roll reuses pandas.Index indexing and append capabilities
  • Index API closely follows DataArray, Dataset and Variable API (i.e., same method names) for consistency
  • Within the Dataset or DataArray method, first call the Index API (if it exists) to create new indexes
    • The Indexes class (i.e., the .xindexes property returns an instance of this class) provides convenient API for iterating through indexes (e.g., get a list of unique indexes, get all coordinates or dimensions for a given index, etc.)
    • If there’s no implementation for the called Index API, either raise an error or fallback to calling the Variable API (below) depending on the case
  • Create new coordinate variables for each of the new indexes using Index.create_variables
    • It is possible to pass a dict of current coordinate variables to Index.create_variables ; it is used to propagate variable metadata (dtype, attrs and encoding)
    • Not all indexes should create new coordinate variables, only those for which it is possible to reuse index data as coordinate variable data (like Pandas indexes)
  • Iterate through the variables and call the Variable API (if it exists)
    • Skip new coordinate variables created at the previous step (just reuse it)
  • Propagate the indexes that are not affected by the operation and clean up all indexes, i.e., ensure consistency between indexes and coordinate variables
    • There is a couple of convenient methods that have been added in #5692 for that purpose: filter_indexes_from_coords and assert_no_index_corrupted
  • Replace indexes and variables, e.g., using _replace, _replace_with_new_dims or _overwrite_indexes methods

Relax all constraints related to “dimension (index) coordinates” in Xarray

Indexes repr

  • Add an Indexes section to Dataset and DataArray reprs
    • #6795
    • #7185
  • Make the repr of Indexes (i.e., .xindexes property) consistent with the repr of Coordinates (.coords property)
  • Add Index._repr_inline_ for tweaking the inline representation of each index shown in the reprs above
    • #7183

Public API for assigning and (re)setting indexes

There is no public API yet for creating and/or assigning existing indexes to Dataset and DataArray objects.

We still need to figure out how best we can (1) assign existing indexes (possibly with their coordinates) and (2) pass index build options.

Other public API for index-based operations

To fully leverage the power and flexibility of custom indexes, we might want to update some parts of Xarray’s public API in order to allow passing arbitrary options per index. For example:

  • sel: the current method and tolerance may not be relevant for all indexes, pass extra arguments to Scipy's cKDTree.query, etc. #7099
  • align: #2217

Also:

  • Make public the Indexes API as it provides convenient methods that might be useful for end-users
  • Import the Index base class into Xarray’s main namespace (i.e., xr.Index)? Also PandasIndex and PandasMultiIndex? The latter may be useful if we depreciate set_index(append=True) and/or if we depreciate “unpacking” pandas.MultiIndex objects to coordinates when given as coords in the Dataset / DataArray constructors.

Documentation

  • User guide:
    • Update the “Terminology” section: “Index” may include custom indexes, review “Dimension coordinate” / “Non-dimension coordinate” as “Indexed coordinate” / “Non-indexed coordinate”
    • Update the “Data structure” section such that it clearly mentions indexes as 1st class citizen of the Xarray data model
    • Maybe update other parts of the documentation that refer to the concept of “dimension coordinate”
  • API reference:
    • add Indexes API
    • add Index API: #6975
  • Xarray internals: add a subsection on how to add custom indexes, maybe with some basic examples: #6975
  • Update development roadmap section

Index types and helper classes built in Xarray

  • Since a lot of potential use-cases for custom indexes may consist in adding some extra logic on top of one or more pandas indexes along one or more dimensions (i.e., “meta-indexes”), it might be worth providing a helper Index abstract subclass that would basically dispatch the given arguments to the corresponding, encapsulated PandasIndex instances and then merge the results
    • #7182
  • Depreciate PandasMultiIndex dimension coordinate?

3rd party indexes

  • Add custom index entrypoint / plugin system, similarly to storage backend entrypoints

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by selecting one unchecked Dataset or DataArray operation from the issue, then read the corresponding Index, PandasIndex, PandasMultiIndex, Indexes, and Variable APIs. Check how the operation creates coordinate variables and preserves or cleans up indexes; it is done when the selected operation supports the intended custom-index behavior and its related tests or documentation are updated.

Written by the indexing model from the issue text.

Assessment

Tech stack
pandas, python
Domain
backend-api-design, data, documentation
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.