make_array: relax element-type equality to accept inputs differing only in nested-field nullability
- Lingua principale
- Rust
- Stelle
- 9.3k
- Fork
- 2.4k
- Merge medio
- 3g 11h
- PR unite (30g)
- 362
Descrizione
## Summary
`make_array` (in `datafusion-functions-nested`) panics when called with arrays whose element types share the same shape but differ in nested-field nullability. Spark, Postgres, and `arrow::compute::concat` all accept this and widen `nullable` to `true` in the result type. DataFusion's `make_array_inner` is stricter, which propagates up to any caller that builds `array(...)` over heterogeneously-produced child expressions.
## Repro symptom
Real-world surfacing in [apache/datafusion-comet](https://github.com/apache/datafusion-comet) on a Delta Lake CDF write that builds `array(struct(id, b, _change_type=lit(\"delete\")), struct(id, b, _change_type=col(...)))` — one arm's `_change_type` is `Utf8` non-nullable (from a literal), another is `Utf8` nullable:
```
panicked at arrow-data-58.2.0/src/transform/mod.rs:422:
assertion `left == right` failed: Arrays with inconsistent types passed to MutableArrayData
left: Struct([Field { name: \"id\", data_type: Int64, nullable: true },
Field { name: \"b\", data_type: Int32 },
Field { name: \"_change_type\", data_type: Utf8 }])
right: Struct([Field { name: \"id\", data_type: Int64, nullable: true },
Field { name: \"b\", data_type: Int32 },
Field { name: \"_change_type\", data_type: Utf8, nullable: true }])
```
Stack: `make_array_inner` → `MutableArrayData::with_capacities`.
## Proposal
`make_array` should accept element types that are equal under nullability-widening (recursively, for nested structs/lists/maps). Concretely:
- Compute the merged element type by walking each child's `DataType` and OR-ing the `nullable` flag at every level (this is essentially `Field::try_merge` minus the type-promotion arm).
- Cast each child to the merged type before handing to `MutableArrayData`.
- Return `ArrayType` with `containsNull = true` if any merge raised a nullability flag.
This matches what `coerce_types`-style coercion does elsewhere in the planner, but applied at execution time when input arrays still disagree (the planner can't always normalize, e.g. when the array is built from disjoint sources like Delta CDF struct literals).
## Why this matters
It blocks native execution of any plan that produces struct elements from multiple sources (CDF writes, UNION ALL inside an `array()`, manually-constructed plans bypassing TypeCoercion). Workaround today: callers must insert explicit casts upstream, or fall back to a non-DataFusion evaluator — both of which lose perf.
## Related caller-side mitigation (for context)
Comet just landed a serde-side decline in [4cb9b4dc](https://github.com/apache/datafusion-comet/commit/) that falls back to Spark's JVM evaluator when `CreateArray`'s children have different `DataType`s. That fix is conservative but loses native execution. Upstreaming the relaxation here would let downstream projects keep native execution and would help any other Arrow-based engine hitting the same shape.
I can put up a PR if the approach lands well.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia nell’implementazione di datafusion-functions-nested, in make_array_inner, e segui la chiamata a MutableArrayData::with_capacities. Confronta la gestione dei tipi esistente con il comportamento descritto in stile coerce_types, quindi verifica che le differenze di nullabilità annidata vengano ampliate, che i figli siano accettati senza un panic e che l’array risultante riporti la nullabilità unificata.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- rust
- Ambito
- data-engineering
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Tranquilla
- Chiarezza
- Specificata chiaramente
- Idoneità per principianti
- 55/100