pydata / pydata/xarray

Opposite of `drop_variables` option in `open_dataset()`

Open
#1,754 3 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

topic-backends
Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

Not so much an issue but more a feature request. xarray.open_dataset() has the drop_variables option, perhaps the opposite could also be useful, e.g. include_variables or read_variables (i.e.; only read the variables which are listed and ignore the others).

I often work with large NetCDF files (>100 variables) out of which I only want to read a few; it is easier to list those few variables than to create a list with the hundred(s) of variables which I don't want.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at xarray.open_dataset() and its existing drop_variables option. Trace how variables are selected when opening NetCDF files, then determine the intended behavior for an include_variables or read_variables option. Done means callers can list a small set of variables and the other variables are ignored.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.