xr.DataArray.sum() converts string objects into unicode
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 4.2k
- Forks
- 1.4k
- Avg merge
- 2d 15h
- Merged PRs (30d)
- 14
Description
What happened:
When summing over all axes of a DataArray with strings of dtype object, the result is a one-size unicode DataArray.
What you expected to happen:
I expected the summation would preserve the dtype, meaning the one-size DataArray would be of dtype object
Minimal Complete Verifiable Example:
ds = xr.DataArray('a', [range(3), range(3)]).astype(object)
ds.sum()
Output
<xarray.DataArray ()>
array('aaaaaaaaa', dtype='<U9')
On the other hand, when summing over one dimension only, the dtype is preserved
ds.sum('dim_0')
Output:
<xarray.DataArray (dim_1: 3)>
array(['aaa', 'aaa', 'aaa'], dtype=object)
Coordinates:
* dim_1 (dim_1) int64 0 1 2
Anything else we need to know?:
The problem becomes relevant as soon as dask is used in the workflow. Dask expects the aggregated DataArray to be of dtype object which will likely lead to errors in the operations to follow.
Probably the behavior comes from creating a new DataArray after the reduction with np.sum() (which itself leads results in a pure python string).
Environment:
Output of xr.show_versions()
INSTALLED VERSIONS
commit: None
python: 3.8.5 (default, Sep 4 2020, 07:30:14)
[GCC 7.3.0]
python-bits: 64
OS: Linux
OS-release: 5.4.0-66-generic
machine: x86_64
processor: x86_64
byteorder: little
LC_ALL: None
LANG: en_US.UTF-8
LOCALE: en_US.UTF-8
libhdf5: 1.10.6
libnetcdf: 4.7.4
xarray: 0.16.2
pandas: 1.2.1
numpy: 1.19.5
scipy: 1.6.0
netCDF4: 1.5.5.1
pydap: None
h5netcdf: 0.7.4
h5py: 3.1.0
Nio: None
zarr: 2.3.2
cftime: 1.3.1
nc_time_axis: None
PseudoNetCDF: None
rasterio: 1.2.0
cfgrib: None
iris: None
bottleneck: 1.3.2
dask: 2021.01.1
distributed: 2021.01.1
matplotlib: 3.3.3
cartopy: 0.18.0
seaborn: 0.11.1
numbagg: None
pint: None
setuptools: 52.0.0.post20210125
pip: 21.0
conda: 4.9.2
pytest: 6.2.2
IPython: 7.19.0
sphinx: 3.4.3
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by running the provided DataArray.sum() example and compare the all-axis reduction with the one-dimension reduction. Trace the DataArray.sum() path through np.sum(), then add a regression test showing that an all-axis sum of object strings retains object dtype and remains compatible with dask.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- numpy, python
- Domain
- data
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 45/100