pydata / pydata/xarray

xr.DataArray.sum() converts string objects into unicode

Open
#5,024 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 15h
Merged PRs (30d)
14

Description

What happened:

When summing over all axes of a DataArray with strings of dtype object, the result is a one-size unicode DataArray.

What you expected to happen:

I expected the summation would preserve the dtype, meaning the one-size DataArray would be of dtype object

Minimal Complete Verifiable Example:

ds = xr.DataArray('a', [range(3), range(3)]).astype(object)
ds.sum()

Output

  <xarray.DataArray ()>
    array('aaaaaaaaa', dtype='<U9')

On the other hand, when summing over one dimension only, the dtype is preserved

ds.sum('dim_0')

Output:

   <xarray.DataArray (dim_1: 3)>
    array(['aaa', 'aaa', 'aaa'], dtype=object)
    Coordinates:
      * dim_1    (dim_1) int64 0 1 2

Anything else we need to know?:

The problem becomes relevant as soon as dask is used in the workflow. Dask expects the aggregated DataArray to be of dtype object which will likely lead to errors in the operations to follow.

Probably the behavior comes from creating a new DataArray after the reduction with np.sum() (which itself leads results in a pure python string).

Environment:

Output of xr.show_versions()

INSTALLED VERSIONS

commit: None
python: 3.8.5 (default, Sep 4 2020, 07:30:14)
[GCC 7.3.0]
python-bits: 64
OS: Linux
OS-release: 5.4.0-66-generic
machine: x86_64
processor: x86_64
byteorder: little
LC_ALL: None
LANG: en_US.UTF-8
LOCALE: en_US.UTF-8
libhdf5: 1.10.6
libnetcdf: 4.7.4

xarray: 0.16.2
pandas: 1.2.1
numpy: 1.19.5
scipy: 1.6.0
netCDF4: 1.5.5.1
pydap: None
h5netcdf: 0.7.4
h5py: 3.1.0
Nio: None
zarr: 2.3.2
cftime: 1.3.1
nc_time_axis: None
PseudoNetCDF: None
rasterio: 1.2.0
cfgrib: None
iris: None
bottleneck: 1.3.2
dask: 2021.01.1
distributed: 2021.01.1
matplotlib: 3.3.3
cartopy: 0.18.0
seaborn: 0.11.1
numbagg: None
pint: None
setuptools: 52.0.0.post20210125
pip: 21.0
conda: 4.9.2
pytest: 6.2.2
IPython: 7.19.0
sphinx: 3.4.3

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running the provided DataArray.sum() example and compare the all-axis reduction with the one-dimension reduction. Trace the DataArray.sum() path through np.sum(), then add a regression test showing that an all-axis sum of object strings retains object dtype and remains compatible with dask.

Written by the indexing model from the issue text.

Assessment

Tech stack
numpy, python
Domain
data
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Clearly specified
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.