duckdb / duckdb/duckdb-python

Python SDK read_parquet union_by_name fail after version 1.2.2

オープン
#259 コメント 5 件 リアクション 0 件 担当者 0 名 GitHub で見る
needs triage
主要言語
Python
スター
187
フォーク
112
平均マージ
13時間 29分
マージ済み PR(30日)
17

説明

[duckdbtest.zip](https://github.com/user-attachments/files/24500181/duckdbtest.zip)

### What happens?

## Title
Regression in 1.3.0+: `union_by_name` fails with "Can't change source type (NULL) to target type (VARCHAR[])" when reading parquet files with mixed NULL/LIST types

## DuckDB Version
- **Working version**: 1.2.2
- **Broken versions**: 1.3.0, 1.3.1 (and later)

## Environment
- OS: Linux
- Python: 3.12.9
- pandas: (latest)

## Description

Starting with DuckDB 1.3.0, reading multiple parquet files with `union_by_name=True` fails when:
1. Some parquet files have a column stored as NULL type (because all values are null in that file)
2. Other parquet files have the same column properly typed as VARCHAR[] (array/list of strings)

This worked correctly in DuckDB 1.2.2 but now throws:
```
BinderException: Binder Error: Can't change source type ("NULL") to target type (VARCHAR[]), type conversion not allowed
```

### Expected Behavior
When `union_by_name=True` is set, DuckDB should merge schemas gracefully, treating NULL-typed columns as compatible with any target type (similar to how pandas handles this).

### Actual Behavior
DuckDB 1.3.0+ throws a `BinderException` and refuses to read the files, even though `union_by_name=True` is explicitly designed to handle schema variations across multiple files.

## Root Cause Analysis

Investigation shows:
- When a parquet file has ALL NULL values for a column, it's stored with NULL type (e.g., `INT32` with `NullType()` logical type)
- Other files with actual data store the same column as `BYTE_ARRAY` with `StringType()` or complex types like `ListType()`
- The error specifically mentions `VARCHAR[]` (array type) suggesting it happens with nested/complex types
- This regression appeared between versions 1.2.2 and 1.3.0

## How to Reproduce

attached files to test see [duckdbtest.zip](https://github.com/user-attachments/files/24500181/duckdbtest.zip)

```python
import duckdb
print(f"DuckDB version: {duckdb.__version__}")

# Fails with 1.3.0+
try:
result = duckdb.read_parquet(
"duckdb_bug_test_files/*.parquet",
union_by_name=True
).df()
print(f"SUCCESS: Read {len(result)} rows")
except Exception as e:
print(f"FAILED: {type(e).__name__}: {e}")
```

### To Reproduce

this is only in python SDK

### OS:

Linux x86

### DuckDB Version:

v1.2.2, v1.3.0 and later

### DuckDB Client:

Python

### Hardware:

_No response_

### Full Name:

Zack Dai

### Affiliation:

Zack Dai

### Did you include all relevant configuration (e.g., CPU architecture, Linux distribution) to reproduce the issue?

- [ ] Yes, I have

### Did you include all code required to reproduce the issue?

- [x] Yes, I have

### Did you include all relevant data sets for reproducing the issue?

Yes

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

まず、添付された duckdbtest.zip を使って、read_parquet に union_by_name=True を指定し、DuckDB 1.2.2 と 1.3.0 以降に対して Python の再現コードを実行します。NULL 列と VARCHAR[] 列のスキーママージ経路を追跡し、リグレッションを特定します。影響を受けるバージョン全体で提供されたファイルが正常に読み込まれ、このケースのリグレッションカバレッジが追加されていれば完了です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
databases
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
静か
明瞭さ
おおむね明確
初心者へのやさしさ
45/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。