aws / aws/amazon-redshift-python-driver

pandas None/NaN mappings

Abierto
#251 1 comentario 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
220
Forks
86
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

I'm using `write_dataframe` function to write a pandas DataFrame. My context is that this dataframe containts columns of three different types:

1. Object (string) in pandas has None as missing data
2. Int64 in pandas has np.nan as missing data
3. Floats64 in pandas has np.nan as missing data

When writing to Redshift, these values are converted as such:

- None as NULL using varchar(20) with bytedict encoding
- NaN as -9223372036854775808 using BIGINT with az64 encoding
- NaN as "NaN" using DOUBLE PRECISION with RAW encoding

When I try to query using SQL, based on the column, I have to filter with:

1. IS NULL
2. = -9223372036854775808
3. ::text = "NaN"

Is this intended? I wish to map all None/NaN values of pandas into NULL values. Is this possible?

Guía de contribución

Abrir la guía de contribución

Línea de trabajo

Start at the write_dataframe entry point and reproduce the issue with object, nullable Int64, and float64 pandas columns containing None or NaN. Trace how each missing value is converted before it reaches Redshift; done means the behavior is consistent with the intended mapping to NULL, with coverage for the three reported column types.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
aws, pandas, python, sql
Área
data, databases
Tipo de issue
Error
Dificultad
3/5
Tiempo estimado
1-2 días
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
35/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.