awslabs / awslabs/python-deequ
[Feature Request]DuckDB as another analytic engine for Deequ
- Vorherrschende Sprache
- Jupyter Notebook
- Sterne
- 826
- Forks
- 158
- Ø Merge
- 9 T. 22 Std.
- Gemergte PRs (30 T.)
- 3
Beschreibung
**Is your feature request related to a problem? Please describe.**
Today, PyDeequ is a PySpark binding for Deequ which is in Scala and Spark only. While it is a good fit for DEs, Spark is not a great fit for many DS use-cases who will have datasets fit in memory and do not want to setup Spark.
See initial ideas here https://youtu.be/fvKFOfaLwBA?t=1393 from @sscdotopen.
**Describe the solution you'd like**
As discussed in above video, it would be good to create the proper abstractions to support another analytic engine. DuckDB who has gained popularity recently can be another analytic engine. The designs need more thoughts/discussion.
Beitragsleitfaden
Rechercherichtung
Beginne damit, die bestehende Beziehung zwischen PyDeequ und PySpark/Scala zu überprüfen, und sieh dir anschließend das verlinkte Video zu den ersten Ideen für Engine-Abstraktionen an. Untersuche, was erforderlich wäre, um DuckDB als alternative Analyse-Engine zu unterstützen. Als erledigt gilt die Aufgabe, wenn das Design abgestimmt ist und ein klarer Implementierungsplan für die Verwendung von Deequ ohne Spark vorliegt.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python, scala
- Bereich
- data-engineering, databases
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Ruhig
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 25/100