apache / apache/datafusion

Substrait plan execution of COUNT() errors with 'pyarrow.lib.ArrowInvalid: Schema at index 0 was different'

Open
#10,873 0 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Rust
Stars
9.3k
Forks
2.4k
Avg merge
3d 7h
Merged PRs (30d)
344

Description

### Describe the bug

Substrait plan execution from COUNT(X) query errors with 'pyarrow.lib.ArrowInvalid: Schema at index 0 was different'

### To Reproduce

```
import json
import pyarrow as pa
import substrait.gen.proto.plan_pb2 as plan_pb2
from datafusion import SessionContext
from datafusion import substrait as ss
from google.protobuf.json_format import Parse
from substrait.gen.proto.plan_pb2 import Plan
from google.protobuf.json_format import MessageToJson

ctx = SessionContext()

tables = pa.RecordBatch.from_arrays(
[
pa.array([1, 2, 3, -4, 5, -6, 7, 8, 9, None]),
],
names=["a"],
)

ctx.register_record_batches("t", [[tables]])

sql_query = "SELECT COUNT(a) FROM 't'"

substrait_proto = plan_pb2.Plan()
substrait_plan = ss.substrait.serde.serialize_to_plan(sql_query, ctx)
substrait_plan_bytes = substrait_plan.encode()
substrait_proto.ParseFromString(substrait_plan_bytes)

substrait_query = MessageToJson(substrait_proto)
substrait_json = json.loads(substrait_query)
plan_proto = Parse(json.dumps(substrait_json), Plan())
plan_bytes = plan_proto.SerializeToString()
substrait_plan = ss.substrait.serde.deserialize_bytes(plan_bytes)
logical_plan = ss.substrait.consumer.from_substrait_plan(ctx, substrait_plan)

df_result = ctx.create_dataframe_from_logical_plan(logical_plan)
df_result.to_arrow_table()
```

Error:
```
Traceback (most recent call last):
File "", line 1, in
File "pyarrow/table.pxi", line 3950, in pyarrow.lib.Table.from_batches
File "pyarrow/error.pxi", line 144, in pyarrow.lib.pyarrow_internal_check_status
File "pyarrow/error.pxi", line 100, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: Schema at index 0 was different:
COUNT(t.a): int64
vs
COUNT(t.a): int64 not null
```

### Expected behavior

_No response_

### Additional context

_No response_

Contributor guide

Open the contributing guide

Research direction

Start with the Python reproduction, especially substrait.serde.deserialize_bytes, substrait.consumer.from_substrait_plan, and df_result.to_arrow_table. Compare the schemas produced before conversion to an Arrow table and verify that COUNT(a) returns consistently nullable metadata without triggering the reported ArrowInvalid error.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, rust, sql
Domain
backend, databases
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.