aws / aws/amazon-redshift-python-driver

pandas None/NaN mappings

未关闭
#251 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
220
派生
86
PR 合并指标
30 天内没有已合并 PR

描述

I'm using `write_dataframe` function to write a pandas DataFrame. My context is that this dataframe containts columns of three different types:

1. Object (string) in pandas has None as missing data
2. Int64 in pandas has np.nan as missing data
3. Floats64 in pandas has np.nan as missing data

When writing to Redshift, these values are converted as such:

- None as NULL using varchar(20) with bytedict encoding
- NaN as -9223372036854775808 using BIGINT with az64 encoding
- NaN as "NaN" using DOUBLE PRECISION with RAW encoding

When I try to query using SQL, based on the column, I have to filter with:

1. IS NULL
2. = -9223372036854775808
3. ::text = "NaN"

Is this intended? I wish to map all None/NaN values of pandas into NULL values. Is this possible?

贡献指南

打开贡献指南

调研方向

Start at the write_dataframe entry point and reproduce the issue with object, nullable Int64, and float64 pandas columns containing None or NaN. Trace how each missing value is converted before it reaches Redshift; done means the behavior is consistent with the intended mapping to NULL, with coverage for the three reported column types.

由索引模型根据 Issue 内容生成。

评估

技术栈
aws, pandas, python, sql
领域
data, databases
Issue 类型
缺陷
难度
3/5
预计耗时
1-2 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。