aws / aws/amazon-redshift-python-driver
The `Cursor.fetch_dataframe` method doesn't respect case-sensitivity
- 主要言語
- Python
- スター
- 220
- フォーク
- 86
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
## Driver version
2.1.3
## Redshift version
PostgreSQL 8.0.2 on i686-pc-linux-gnu, compiled by GCC gcc (GCC) 3.4.2 20041017 (Red Hat 3.4.2-6.fc3), Redshift 1.0.76169
## Client Operating System
macos
## Python version
3.12.6
## Problem description
The `Cursor.fetch_dataframe` method [always lowercases](https://github.com/aws/amazon-redshift-python-driver/blob/64cbd54ef6a5e71d98ce193630d09267cc154379/redshift_connector/cursor.py#L526) column names, irrespective of the [case-sensitivity configuration value](https://docs.aws.amazon.com/redshift/latest/dg/r_enable_case_sensitive_identifier.html). This behavior is unexpected, because the resulting `DataFrame`'s columns may be treated as case-insensitive, even when that flag is set to `true`.
One example where this can be problematic is demonstrated below:
```python
from redshift_connector import connect
conn = connect()
cursor = conn.cursor()
cursor.execute('SET enable_case_sensitive_identifier TO true')
cursor.execute('WITH t AS (SELECT 1 AS "C", 2 AS "c") SELECT * FROM t')
# cursor.fetch_dataframe()
# c c
# 0 1 2
# cursor.fetch_dataframe().to_dict()
# :1: UserWarning: DataFrame columns are not unique, some columns will be omitted.
# {'c': {0: 2}}
```
## Possible solutions
I see that there is an [open PR](https://github.com/aws/amazon-redshift-python-driver/pull/232) related to this issue, but I don't think it solves it. I believe that the correct way to solve this is to get rid of the `lower()` call in line 526 altogether. That would mean that the columns produced by the driver reflect those returned by Redshift, hence respecting the case-sensitivity configuration value (see image below). I plan to open a PR with this fix soon.
----
P.S.: I see that the line of interest was introduced in the [first commit](https://github.com/aws/amazon-redshift-python-driver/blame/64cbd54ef6a5e71d98ce193630d09267cc154379/redshift_connector/cursor.py#L526) of this repo (!) and hasn't changed since. My hunch is that this was most likely an oversight, since Redshift's [documentation](http://web.archive.org/web/20200804231541/https://docs.aws.amazon.com/redshift/latest/dg/welcome.html) at the date of that commit had no mention of the `enable_case_sensitive_identifier` flag, so the driver must've not been updated to take it into account after it was introduced.
コントリビューションガイド
調査の方向性
redshift_connector/cursor.py の 526 行目付近にある fetch_dataframe の実装から始めます。ここでは列名が小文字に変換されています。enable_case_sensitive_identifier を true に設定し、"C" と "c" の両方を返すクエリを使って問題を再現します。完了条件は、結果の DataFrame が、大小文字の違いを含め、Redshift から返された列名を保持することです。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- postgresql, python
- 領域
- databases
- issue の種類
- バグ
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 活発さ
- 停滞
- 明瞭さ
- 明確に書かれている
- 初心者へのやさしさ
- 50/100