DB engine for pandas: sql.connect or sqlalchemy

未关闭
#476 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
25/100
Issue 类型
文档
描述清晰度
需要澄清
活跃度
停滞
技术栈
pandas, python, sql, sqlalchemy
领域
data, databases

调研方向

该 issue 比较 databricks.sql.connect 与 SQLAlchemy 和 pandas.read_sql,包括 CloudFetch 以及可能的序列化开销。首先查看链接的、与 pandas 相关的已关闭 pull request,以及 pandas.read_sql 集成。完成的标准是记录推荐的 API、性能影响,以及两条路径是否都保留 CloudFetch。

由索引模型根据 Issue 内容生成。

描述

Hello,

I was wondering what's the best practice for using this package with pandas.

  1. It's possible to create a databricks.sql.connect and pass it to pandas.read_sql. This works however it raises
UserWarning: pandas only supports SQLAlchemy connectable (engine/connection) or database string URI or sqlite3 DBAPI2 
connection. Other DBAPI2 objects are not tested. Please consider using SQLAlchemy.
  1. Alternatively it's possible to use SQLAlchemy with a databricks:// URL and pass that to pandas. Doesn't it mean an extra serialization step performance wise though?

What's the recommended way, in particular regarding performance? Would both use CloudFetch for larger queries? I see there are some fixes/improvements done for pandas done in PRs so which API should be used to benefit from those?

Thanks!

cc @kravets-levko

主要语言
Python
星标
233
派生
152
平均合并
21 小时 5 分钟
30 天内合并 PR
10

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

databricks/databricks-sql-python 的其他 Issue

查看 databricks/databricks-sql-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。