airbytehq / airbytehq/PyAirbyte
💡 Feature Request: Add `to_spark()` to on dataset class
未关闭
1.0
accepting pull requests
good first issue
spark
- 主要语言
- Python
- 星标
- 344
- 派生
- 77
- 平均合并
- 1 天 11 小时
- 30 天内合并 PR
- 35
描述
There's an interesting use case here where you could use PyAirbyte directly in data pipelines that run on spark.
Currently if you want to do this, you need to do `to_pandas()` and then `spark_session.createDataFrame(issues_df, shema=my_schema)`, but this seems inefficient, plus you have to manually define the schema (for example for json blobs which are `object` in pandas but need to be `StringType` in spark, and other idiosyncrasies like pandas having 64 bit ints but spark having Int and Long).
Or maybe a spark df cache would be more efficient here?
贡献指南
评估
这个 Issue 还没有评估数据。