apache / apache/arrow

[Python] `is_null(nan_is_null=True)` does not work with only NaN's

Open
#34,162 3 comments 0 reactions 0 assignees View on GitHub
Component: Python Type: bug
Dominant language
C++
Stars
17.1k
Forks
4.3k
Avg merge
3d 18h
Merged PRs (30d)
91

Description

### Describe the bug, including details regarding any error messages, version, and platform.

I was working on some test-cases for the PyIceberg integration, and hit this edge case. When you have a file with only NaN values, it will be skipped when reading the file with a `is_null(nan_is_null=True)` filter.

In Spark, I create the following table:

```sql
CREATE TABLE test_null_nan
USING iceberg
AS SELECT
1 AS idx,
float('NaN') AS col_numeric
UNION ALL SELECT
2 AS idx,
null AS col_numeric
UNION ALL SELECT
3 AS idx,
1 AS col_numeric
```

This then creates three files with each one record:

```
➜ python git:(fd-integration-tests) ✗ pyiceberg --catalog local files default.test_null_nan
Snapshots: local.default.test_null_nan
└── Snapshot 870844541941792785, schema 0: s3a://warehouse/wh/default/test_null_nan/metadata/snap-870844541941792785-1-a05e1621-f735-4837-bb86-ce9886da3e6b.avro
└── Manifest: s3a://warehouse/wh/default/test_null_nan/metadata/a05e1621-f735-4837-bb86-ce9886da3e6b-m0.avro
├── Datafile: s3a://warehouse/wh/default/test_null_nan/data/00000-0-658408d0-d063-4caa-b310-f68552713bea-00001.parquet
├── Datafile: s3a://warehouse/wh/default/test_null_nan/data/00001-1-5e625fcb-4a0c-4082-9371-7f4897768ccd-00001.parquet
└── Datafile: s3a://warehouse/wh/default/test_null_nan/data/00002-2-11de56ee-27c1-45ff-be61-cc52727c1b84-00001.parquet
```

If I filter using `pc.col('col_numeric').is_null(nan_is_null=True) & ~pc.col('col_numeric').is_null()` I don't get any results. When I rewrite the table into a single file:

```sql
CREATE TABLE test_null_nan_rewritten
USING iceberg
AS SELECT * FROM test_null_nan
```

And then do the same filter operation, I do get results. I suspect there is something off with the page skipping when `nan_is_null=True`.

### Component(s)

Python

Contributor guide

Open the contributing guide

Research direction

Start at the Python `pc.col(...).is_null` entry point and the page-skipping path implicated by the report. Reproduce the filter against the three-file and rewritten single-file cases, then verify that the NaN-only file is handled consistently and the filter returns the expected results.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.