airbytehq / airbytehq/PyAirbyte

💡 Feature Request: Add `to_spark()` to on dataset class

Aperta
#173 5 commenti 0 reazioni 0 assegnatari Vedi su GitHub
1.0 accepting pull requests good first issue spark
Lingua principale
Python
Stelle
344
Fork
77
Merge medio
1g 11h
PR unite (30g)
35

Descrizione

There's an interesting use case here where you could use PyAirbyte directly in data pipelines that run on spark.

Currently if you want to do this, you need to do `to_pandas()` and then `spark_session.createDataFrame(issues_df, shema=my_schema)`, but this seems inefficient, plus you have to manually define the schema (for example for json blobs which are `object` in pandas but need to be `StringType` in spark, and other idiosyncrasies like pandas having 64 bit ints but spark having Int and Long).

Or maybe a spark df cache would be more efficient here?

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.