pydata / pydata/xarray

tolerance for alignment

Open
#2,217 23 comments 11 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

topic-combine
Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

When using open_mfdataset on files which 'should' share a grid, there is often a small mismatch which results in the grid not aligning properly. This happens frequently when trying to read data from large climate models from multiple files of the same variable, same lon,lat grid and different time intervals. This silent behavior means that I always have to check the sizes of the lon,lat grids whenever I rely on mfdataset to concatenate the data in time.

Here is an example in which I create two 1d DataArrays which have slightly different coordinates:

import xarray as xr
import numpy as np
from glob import glob

tol=1e-14
x1 = np.arange(1,6)+ tol*np.random.rand(5)
da1 = xr.DataArray([9, 0, 2, 1, 0], dims=['x'], coords={'x': x1})

x2 = np.arange(1,6) + tol*np.random.rand(5)
da2 = da1.copy()
da2['x'] = x2

print(da1.x,'\n', da2.x)
<xarray.DataArray 'x' (x: 5)>
array([1., 2., 3., 4., 5.])
Coordinates:
  * x        (x) float64 1.0 2.0 3.0 4.0 5.0 
 <xarray.DataArray 'x' (x: 5)>
array([1., 2., 3., 4., 5.])
Coordinates:
  * x        (x) float64 1.0 2.0 3.0 4.0 5.0

First I save both DataArrays as netcdf files and then use open_mfdataset to load them:

da1.to_netcdf('da1.nc',encoding={'x':{'dtype':'float64'}})
da2.to_netcdf('da2.nc',encoding={'x':{'dtype':'float64'}})

db = xr.open_mfdataset(glob('da?.nc'))

db
<xarray.Dataset>
Dimensions:                        (x: 10)
Coordinates:
  * x                              (x) float64 1.0 2.0 3.0 4.0 5.0 1.0 2.0 ...
Data variables:
    __xarray_dataarray_variable__  (x) int64 dask.array<shape=(10,), chunksize=(5,)>

So the x grid is now twice the size. This behavior is the same if I just use align with join='outer':

xr.align(da1,da2,join='outer')
(<xarray.DataArray (x: 10)>
 array([nan,  9., nan,  0.,  2., nan, nan,  1.,  0., nan])
 Coordinates:
   * x        (x) float64 1.0 1.0 2.0 2.0 3.0 3.0 4.0 4.0 5.0 5.0,
 <xarray.DataArray (x: 10)>
 array([ 9., nan,  0., nan, nan,  2.,  1., nan, nan,  0.])
 Coordinates:
   * x        (x) float64 1.0 1.0 2.0 2.0 3.0 3.0 4.0 4.0 5.0 5.0)
Request/ suggestion

What is needed is a user specified tolerance level to give to open_mfdataset and passed to
align which will accept these grids as the same

Possibly related to https://github.com/pydata/xarray/issues/2215

xr.__version__ '0.10.4'

thanks, Naomi

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the example using the shown DataArrays, netCDF files, open_mfdataset, and xr.align(join='outer'). Read the requested tolerance behavior and the related issue 2215 before defining the scope. Done means a user-specified tolerance can be passed through open_mfdataset to alignment so nearly identical grids are treated as the same.

Written by the indexing model from the issue text.

Assessment

Tech stack
numpy, python
Domain
data
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.