pydata / pydata/xarray

Add `nunique` reduction for number of unique values

Open
#9,548 7 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

contrib-good-first-issue enhancement topic-pandas-like
Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

Is your feature request related to a problem?

From https://github.com/pydata/xarray/issues/9544#issuecomment-2372685411

Though perhaps we should add nunique along a dimension implemented as sort along axis, succeeding-elements-are-not-equal along axis handling NaNs, then sum along axis.

xref pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.nunique.html

I think I'd add it to https://github.com/pydata/xarray/blob/main/xarray/util/generate_aggregations.py

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading xarray/util/generate_aggregations.py and the existing reduction definitions it generates. Confirm how reductions handle dimensions and NaNs, then add the nunique reduction following the issue's described behavior. Done means xarray exposes a number-of-unique-values reduction along a dimension.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.