apache / apache/iceberg-python
Error when upserting and updating dataframes with dictionary encoded columns
- 主要言語
- Python
- スター
- 1.1k
- フォーク
- 581
- 平均マージ
- 1日 17時間
- マージ済み PR(30日)
- 77
説明
### Apache Iceberg version
0.9.1
### Please describe the bug 🐞
df = pyarrow.read_table(read_dictionary=[string_columns])
table = catalog.load_table(table_name)
table.append(df) -> 'append' casts dictionary columns into large string and table appended without any error
table.upsert(df, join_cols=[primary_keys]) -> Error (Invalid Type Dictionary)
table.update(df, join_cols=[primary_keys]) -> Error (Invalid Type Dictionary)
### Willingness to contribute
- [ ] I can contribute a fix for this bug independently
- [ ] I would be willing to contribute a fix for this bug with guidance from the Iceberg community
- [x] I cannot contribute a fix for this bug at this time
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
調査の方向性
まず、pyarrow.read_table(read_dictionary=[string_columns]) と catalog.load_table(table_name) から読み込んだテーブルで問題を再現します。table.append(df)、table.upsert(df, join_cols=[primary_keys])、table.update(df, join_cols=[primary_keys]) を比較します。upsert と update が、報告されている Invalid Type Dictionary エラーなしで dictionary エンコードされた列を処理できれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- databases
- issue の種類
- バグ
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 活発さ
- 活発
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 50/100