duckdb / duckdb/duckdb-python

Columns names outputs are inconsistents and case sensitive

Open
#418 0 comments 0 reactions 0 assignees View on GitHub
needs triage
Dominant language
Python
Stars
186
Forks
113
Avg merge
13h 29m
Merged PRs (30d)
17

Description

### What happens?

Output of the code block:
```shell
PS C:\Users\tibo\python_codes\pql> uv run t.py
relation.columns: ['foo', 'Foo']
to_arrow_table: ['foo', 'Foo']
pl.from_arrow: ['foo', 'Foo']
relation.pl eager: ['foo', 'Foo_1']
relation.pl lazy: ['foo', 'Foo']
pandas result: ['foo', 'Foo_1']
arrow query result: ['foo', 'Foo']
PS C:\Users\tibo\python_codes\pql>
```

pandas and polars dataframe don't give the same results as the rest.

### To Reproduce

```python
import duckdb
import polars as pl

rel = duckdb.from_query("select 1 as foo, 2 as Foo")

print("relation.columns:", rel.columns)
print("to_arrow_table:", rel.to_arrow_table().column_names)
print("pl.from_arrow:", pl.from_arrow(rel.to_arrow_table()).columns)
print("relation.pl eager:", rel.pl().columns)
print("relation.pl lazy:", rel.pl(lazy=True).collect().columns)
print("pandas result:", list(rel.df().columns))
print("arrow query result:", rel.to_arrow_reader().read_all().column_names)
```

### OS:

Windows

### DuckDB Package Version:

1.5.1

### Python Version:

3.13.7

### Full Name:

Stettler Thibaud

### Affiliation:

University of Geneva

### What is the latest build you tested with? If possible, we recommend testing with the latest nightly build.

`1.5.2.dev40`

### Did you include all relevant data sets for reproducing the issue?

Yes

### Did you include all code required to reproduce the issue?

- [x] Yes, I have

### Did you include all relevant configuration to reproduce the issue?

- [x] Yes, I have

Contributor guide

Open the contributing guide

Research direction

Start by running the supplied Python reproduction against the current DuckDB Python package and compare rel.pl(), rel.df(), the Arrow conversions, and the lazy result. Trace the Python relation-to-Polars and relation-to-pandas entry points, then add or update regression coverage so equivalent outputs handle the case-variant columns consistently.

Written by the indexing model from the issue text.

Assessment

Tech stack
pandas, python
Domain
data, database
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.