Arrow scan: a NULL row passes a pushed-down `<=` filter against a NaN constant
- Vorherrschende Sprache
- Python
- Sterne
- 187
- Forks
- 112
- Ø Merge
- 13 Std. 29 Min.
- Gemergte PRs (30 T.)
- 17
Beschreibung
### What happens?
When a `WHERE b <= 'nan'::DOUBLE` filter is pushed into an Arrow scan, the scan returns rows where b is NULL.
The same predicate over same rows copied to a native table (correctly) excludes the null rows.
This *should* exclude the nulls (`NULL <= x` is false): The bug appears to be in the arrow filter_pushdown: disabling the optimizer/pushdowns returns the correct results.
### To Reproduce
```py
import duckdb
import pyarrow as pa
with duckdb.connect() as con:
print(f"duckdb {duckdb.__version__}, pyarrow {pa.__version__}", "source id:", con.execute("SELECT source_id FROM pragma_version()").fetchone()[0])
arrow_table = pa.table({"b": pa.array([1.0, float("nan"), None])})
con.register("arrow_t", arrow_table)
result_native = con.execute("create table native_t as select * from arrow_t;select b from native_t where b <= 'nan'::DOUBLE").fetchall()
result = con.execute("select b from arrow_t where b <= 'nan'::DOUBLE").fetchall()
con.execute("PRAGMA disable_optimizer;")
result_no_optimizer = con.execute("select b from arrow_t where b <= 'nan'::DOUBLE").fetchall()
print(f"{result_native=}")
print(f"{result=}")
print(f"{result_no_optimizer=}")
```
## Result: Note that "result" differs from others
>
> duckdb 1.6.0.dev379, pyarrow 25.0.1 source id: a00803f768
> result_native=[(1.0,), (nan,)]
> result=[(1.0,), (nan,), (None,)]
> result_no_optimizer=[(1.0,), (nan,)]
### OS:
Windows & Ubuntu WSL
### DuckDB Package Version:
1.5.5 and 1.6.0.dev379
### Python Version:
3.14
### Full Name:
Paul T
### Affiliation:
Iqmo
### What is the latest build you tested with? If possible, we recommend testing with the latest nightly build.
I have tested with a nightly build
### Did you include all relevant data sets for reproducing the issue?
Yes
### Did you include all code required to reproduce the issue?
- [x] Yes, I have
### Did you include all relevant configuration to reproduce the issue?
- [x] Yes, I have
Beitragsleitfaden
Rechercherichtung
Beginne mit der Python-Reproduktion unter Verwendung der registrierten PyArrow-Tabelle und vergleiche den Arrow-Scan mit den Ergebnissen von native-table und optimizer-disabled. Verfolge das Arrow-Filter-Pushdown für `b <= 'nan'::DOUBLE`; als erledigt gilt die Aufgabe, wenn die per Pushdown ausgeführte Abfrage die NULL-Zeile wie die anderen Ergebnisse ausschließt und ein Regressionstest diesen Fall abdeckt.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- databases
- Issue-Typ
- Bug
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- 1-2 Tage
- Aktivitätsstatus
- Aktiv
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 68/100