Python SDK read_parquet union_by_name fail after version 1.2.2
- 主要言語
- Python
- スター
- 187
- フォーク
- 112
- 平均マージ
- 13時間 29分
- マージ済み PR(30日)
- 17
説明
[duckdbtest.zip](https://github.com/user-attachments/files/24500181/duckdbtest.zip)
### What happens?
## Title
Regression in 1.3.0+: `union_by_name` fails with "Can't change source type (NULL) to target type (VARCHAR[])" when reading parquet files with mixed NULL/LIST types
## DuckDB Version
- **Working version**: 1.2.2
- **Broken versions**: 1.3.0, 1.3.1 (and later)
## Environment
- OS: Linux
- Python: 3.12.9
- pandas: (latest)
## Description
Starting with DuckDB 1.3.0, reading multiple parquet files with `union_by_name=True` fails when:
1. Some parquet files have a column stored as NULL type (because all values are null in that file)
2. Other parquet files have the same column properly typed as VARCHAR[] (array/list of strings)
This worked correctly in DuckDB 1.2.2 but now throws:
```
BinderException: Binder Error: Can't change source type ("NULL") to target type (VARCHAR[]), type conversion not allowed
```
### Expected Behavior
When `union_by_name=True` is set, DuckDB should merge schemas gracefully, treating NULL-typed columns as compatible with any target type (similar to how pandas handles this).
### Actual Behavior
DuckDB 1.3.0+ throws a `BinderException` and refuses to read the files, even though `union_by_name=True` is explicitly designed to handle schema variations across multiple files.
## Root Cause Analysis
Investigation shows:
- When a parquet file has ALL NULL values for a column, it's stored with NULL type (e.g., `INT32` with `NullType()` logical type)
- Other files with actual data store the same column as `BYTE_ARRAY` with `StringType()` or complex types like `ListType()`
- The error specifically mentions `VARCHAR[]` (array type) suggesting it happens with nested/complex types
- This regression appeared between versions 1.2.2 and 1.3.0
## How to Reproduce
attached files to test see [duckdbtest.zip](https://github.com/user-attachments/files/24500181/duckdbtest.zip)
```python
import duckdb
print(f"DuckDB version: {duckdb.__version__}")
# Fails with 1.3.0+
try:
result = duckdb.read_parquet(
"duckdb_bug_test_files/*.parquet",
union_by_name=True
).df()
print(f"SUCCESS: Read {len(result)} rows")
except Exception as e:
print(f"FAILED: {type(e).__name__}: {e}")
```
### To Reproduce
this is only in python SDK
### OS:
Linux x86
### DuckDB Version:
v1.2.2, v1.3.0 and later
### DuckDB Client:
Python
### Hardware:
_No response_
### Full Name:
Zack Dai
### Affiliation:
Zack Dai
### Did you include all relevant configuration (e.g., CPU architecture, Linux distribution) to reproduce the issue?
- [ ] Yes, I have
### Did you include all code required to reproduce the issue?
- [x] Yes, I have
### Did you include all relevant data sets for reproducing the issue?
Yes
コントリビューションガイド
調査の方向性
まず、添付された duckdbtest.zip を使って、read_parquet に union_by_name=True を指定し、DuckDB 1.2.2 と 1.3.0 以降に対して Python の再現コードを実行します。NULL 列と VARCHAR[] 列のスキーママージ経路を追跡し、リグレッションを特定します。影響を受けるバージョン全体で提供されたファイルが正常に読み込まれ、このケースのリグレッションカバレッジが追加されていれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- databases
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 静か
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 45/100