apache / apache/arrow-julia

Dense Union incompatible between Julia/Python

Open
#285 2 comments 1 reaction 0 assignees View on GitHub
Dominant language
Julia
Stars
312
Forks
78
PR merge metrics
No merged PRs in 30d

Description

When writing a table with Arrow.jl that contains a nullable column, the Arrow data cannot be read by Pyarrow:

```julia
julia> Arrow.write("/tmp/nothing.arrow", (col=Vector{Union{Nothing,Int32}}([1,2,3,nothing]),))
"/tmp/nothing.arrow"

julia> Arrow.Table("/tmp/nothing.arrow")
Arrow.Table with 4 rows, 1 columns, and schema:
:col Union{Missing, Nothing, Int32}
```

```python
In [9]: df = pandas.read_feather("/tmp/nothing.arrow")
-----------------------------------------------------------------
ArrowNotImplementedError Traceback (most recent call last)
in
----> 1 df = pandas.read_feather("/tmp/nothing.arrow")

~/miniconda3/lib/python3.9/site-packages/pandas/io/feather_format.py in read_feather(path, columns, use_threads, storage_options)
128 ) as handles:
129
--> 130 return feather.read_feather(
131 handles.handle, columns=columns, use_threads=bool(use_threads)
132 )

~/miniconda3/lib/python3.9/site-packages/pyarrow/feather.py in read_feather(source, columns, use_threads, memory_map)
218 """
219 _check_pandas_version()
--> 220 return (read_table(source, columns=columns, memory_map=memory_map)
221 .to_pandas(use_threads=use_threads))
222

~/miniconda3/lib/python3.9/site-packages/pyarrow/array.pxi in pyarrow.lib._PandasConvertible.to_pandas()

~/miniconda3/lib/python3.9/site-packages/pyarrow/table.pxi in pyarrow.lib.Table._to_pandas()

~/miniconda3/lib/python3.9/site-packages/pyarrow/pandas_compat.py in table_to_blockmanager(options, table, categories, ignore_metadata, types_mapper)
787 _check_data_column_metadata_consistency(all_columns)
788 columns = _deserialize_column_index(table, all_columns, column_indexes)
--> 789 blocks = _table_to_blocks(options, table, categories, ext_columns_dtypes)
790
791 axes = [columns, index]

~/miniconda3/lib/python3.9/site-packages/pyarrow/pandas_compat.py in _table_to_blocks(options, block_table, categories, extension_columns)
1126 # Convert an arrow table to Block from the internal pandas API
1127 columns = block_table.column_names
-> 1128 result = pa.lib.table_to_blocks(options, block_table, categories,
1129 list(extension_columns.keys()))
1130 return [_reconstruct_block(item, columns, extension_columns)

~/miniconda3/lib/python3.9/site-packages/pyarrow/table.pxi in pyarrow.lib.table_to_blocks()

~/miniconda3/lib/python3.9/site-packages/pyarrow/error.pxi in pyarrow.lib.check_status()

ArrowNotImplementedError: No known equivalent Pandas block for Arrow data of type dense_union<: null=0, : int32 not null=1> is known.
```

Note that when using Missing instead of Nothing Pyarrow can read the data written by Arrow.jl.

```julia
julia> Arrow.write("/tmp/missing.arrow", (col=Vector{Union{Missing,Int32}}([1,2,3,missing]),))
"/tmp/missing.arrow"
```

```python
In [1]: pandas.read_feather("/tmp/missing.arrow")
Out[1]:
col
0 1.0
1 2.0
2 3.0
3 NaN
```

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the Julia Arrow.write example with Vector{Union{Nothing,Int32}} and compare it with the working Missing variant, then run pandas.read_feather on both files. Trace how Arrow.Table represents the nullable column and how PyArrow handles the resulting dense_union; done means the Nothing-written file can be read by PyArrow without the reported ArrowNotImplementedError.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.