airbytehq / airbytehq/PyAirbyte
💡 Feature Request: Add `to_spark()` to on dataset class
Đang mở
1.0
accepting pull requests
good first issue
spark
- Ngôn ngữ chính
- Python
- Star
- 344
- Fork
- 77
- Merge trung bình
- 1 ngày 11 giờ
- Pull request đã merge (30 ngày)
- 35
Mô tả
There's an interesting use case here where you could use PyAirbyte directly in data pipelines that run on spark.
Currently if you want to do this, you need to do `to_pandas()` and then `spark_session.createDataFrame(issues_df, shema=my_schema)`, but this seems inefficient, plus you have to manually define the schema (for example for json blobs which are `object` in pandas but need to be `StringType` in spark, and other idiosyncrasies like pandas having 64 bit ints but spark having Int and Long).
Or maybe a spark df cache would be more efficient here?
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.