awslabs / awslabs/python-deequ

[Pydeequ 1.0.1] pydeequ.checks.isContainedIn does not accept lambda assertion

Offen
#88 4 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
bug
Vorherrschende Sprache
Jupyter Notebook
Sterne
826
Forks
158
Ø Merge
9 T. 22 Std.
Gemergte PRs (30 T.)
3

Beschreibung

**Describe the bug**
When running Pydeequ 1.0.1 the test generated by ConstraintSuggestionRunner include tests using the isContainedIn() function that fail during execution.

The cause is that the suggested tests include a lambda assertion which the python function does not accept as it takes 3 positional arguments but the suggested tests has 5

**To Reproduce**
Steps to reproduce the behavior:
1. Generate a test statement using a dataset that is incomplete, resulting in a suggestion for a test using isContainedIn() which uses a lambda:

- Example: 'value'is empty for more then 97% of the records:
- isContainedIn("value", [""], lambda x: x >= 0.97, "It should be above 0.97!")

2. Execute the test
3. Check output for error:
- TypeError: isContainedIn() takes 3 positional arguments but 5 were given

This issue has been reported before: https://github.com/awslabs/python-deequ/issues/65

The cause is that the current implementation of the isContainedIn was edited in https://github.com/awslabs/python-deequ/commit/30375bb8645728a539b7b2f6d2d85f89266ac047#diff-783716851e9837b9753e643de1f15e031f79bed4ef27e07ce67eeddc5a3fb2ee but the ConstraintSuggestionRunner was not updated to match the latest implementation.

It is unclear to me whether the suggested test is valid and the isContainedIn function needs to be extended or whether the change was made for a reason and thus the ConstraintSuggestionRunner should be adjusted to leave out the broken tests.

A previously made pull request does show how to revert the change: https://github.com/awslabs/python-deequ/pull/58

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Beginne mit der aktuellen Implementierung von isContainedIn und ConstraintSuggestionRunner und reproduziere dann das Problem mit einem unvollständigen Datensatz, der den auf lambda basierenden Vorschlag erzeugt. Vergleiche die Aufrufsignatur mit dem generierten Test und verifiziere, dass der Vorschlag ohne den gemeldeten TypeError ausgeführt wird.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
testing
Issue-Typ
Bug
Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Aktivitätsstatus
Ruhig
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
45/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.