apache / apache/arrow

[Python] Pickling a sliced array serializes all the buffers

Open
#26,685 37 comments 1 reaction 1 assignee Claimed by @anjakefala View on GitHub
Component: Python Priority: Critical Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 13h
Merged PRs (30d)
88

Description

If a large array is sliced, and pickled, it seems the full buffer is serialized, this leads to excessive memory usage and data transfer when using multiprocessing or dask.
```java

>>> import pyarrow as pa
>>> ar = pa.array(['foo'] * 100_000)
>>> ar.nbytes
700004
>>> import pickle
>>> len(pickle.dumps(ar.slice(10, 1)))
700165

NumPy for instance
>>> import numpy as np
>>> ar_np = np.array(ar)
>>> ar_np
array(['foo', 'foo', 'foo', ..., 'foo', 'foo', 'foo'], dtype=object)
>>> import pickle
>>> len(pickle.dumps(ar_np[10:11]))
165
```
I think this makes sense if you know arrow, but kind of unexpected as a user.

Is there a workaround for this? For instance copy an arrow array to get rid of the offset, and trim the buffers?

**Reporter**: [Maarten Breddels](https://issues.apache.org/jira/browse/ARROW-10739) / @maartenbreddels
**Assignee**: [Clark Zinzow](https://issues.apache.org/jira/browse/ARROW-10739)

#### Related issues:
- https://github.com/apache/arrow/issues/30503 (is related to)

**Note**: *This issue was originally created as [ARROW-10739](https://issues.apache.org/jira/browse/ARROW-10739). Please see the [migration documentation](https://github.com/apache/arrow/issues/14542) for further details.*

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.