awslabs / awslabs/deequ

[BUG] Row-level filtering marking the records as pass when null values are present in the column

Open
#565 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Scala
Stars
3.6k
Forks
586
Avg merge
13d 13h
Merged PRs (30d)
1

Description

I am working on filtering data based on row-level checks. It's working fine when notnull values are present in the column

But incorrectly marking the records as pass when null values are present in the column.

For example
```
import sparkSession.implicits._

Seq(
(1, "a", 1),
(2, "b", 3),
(3, null, null),
(4, "c", 5),
(5, null, null),
(6, "d", 7)
).toDF("item", "att1", "att2")
```

Applied below rules:

```
rule1 : .isPrimaryKey("att1","att2")
rule2: .isGreaterThan("att2", "att1")
rule3: .isgreaterthanorequalto("att2", "att1")
+----+----+----+-----+-----+-----+
|item|att1|att2|rule1|rule2|rule3|
+----+----+----+-----+-----+-----+
| 1| a| 1| true|false| true|
| 2| b| 3| true| true| true|
| 3|null|null| true| true| true|
| 4| c| 5| true| true| true|
| 5|null|null| true| true| true|
| 6| d| 7| true| true| true|
+----+----+----+-----+-----+-----+
```

When columns values are null, the row-level check status is considered as true but it should be false.

Contributor guide

Open the contributing guide

Research direction

Reproduce the reported Spark DataFrame and row-level checks using isPrimaryKey, isGreaterThan, and isGreaterThanOrEqualTo. Trace the filtering path and compare how null-valued rows are evaluated; done means those rows receive false status while existing non-null behavior remains unchanged.

Written by the indexing model from the issue text.

Assessment

Tech stack
scala, spark
Domain
data, testing
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.