pydata / pydata/xarray

nc file locked by xarray after (double) da.compute() call

Open
#3,041 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

topic-backends topic-dask
Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

Hi,
I've recently started to use dask.array to compute and load only chunks of a huge dataset for the development of Parcels. I've run into an issue of locked netcdf files while running multiple tests using py.test, and I think my bug should be related to the reduced issue that I reproduce hereinafter:

Problem Description
import numpy as np
import xarray as xr
import dask.array as da

def create_ds():
    temp = 15 + 8 * np.random.randn(2, 2)
    precip = 10 * np.random.rand(2, 2)
    lon = [[-99.83, -99.32], [-99.79, -99.23]]
    lat = [[42.25, 42.21], [42.63, 42.59]]
    ds = xr.Dataset({'temperature': (['x', 'y'],  temp),
                     'precipitation': (['x', 'y'], precip)},
                      coords={'lon': (['x', 'y'], lon),
                              'lat': (['x', 'y'], lat)})
    ds.to_netcdf('test.nc')

def dask_op():
    ds = xr.open_dataset('test.nc')
    temp = da.from_array(ds.temperature, chunks='auto')
    temp.compute()
    temp.compute()

create_ds()
dask_op()
dset = xr.Dataset()
dset.to_netcdf('test.nc')
Output

The last line of the mini-code crashes since test.nc is still locked.
Of course the problem can be circumvented by closing the dataset after last call to temp.compute, but this operation is not necessary if temp.compute() is called only once?

Working and not working alternatives

If I replace the 2 temp.compute() lines by either:

temp.compute()

or

temp.compute()
temp.compute()
ds.close

The code passes, but with those blocks:

temp.compute()
temp.compute()

or

ds.close
temp.compute()
temp.compute()
ds.close

it doesn't.

This issue was also posted on dask, where I was advised to also ask for help here.
Any hint about the reason for this issue? Thanks!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the reproducer in the issue with create_ds() and dask_op(), comparing one and two temp.compute() calls against explicit ds.close() calls. Investigate how the opened test.nc dataset remains locked, and consider the issue resolved when the double-compute case no longer leaves test.nc locked without requiring an unnecessary close.

Written by the indexing model from the issue text.

Assessment

Tech stack
numpy, python
Domain
data
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.