apache / apache/iceberg-python

Avro schema is re-converted to an Iceberg schema on every manifest read during scan planning

未关闭 适合新手
#3,662 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
1.1k
派生
581
平均合并
1 天 17 小时
30 天内合并 PR
78

描述

### Feature Request / Improvement

`AvroFileHeader.get_schema()` (`pyiceberg/avro/file.py`) runs the full `avro_to_iceberg` conversion every time an
Avro file is opened.

**The inefficiency**

- Every manifest under a spec embeds an *identical* Avro schema string.
- So during scan planning, that same conversion is repeated **once per manifest**.
- The cost grows with the manifest count, even though only a couple of distinct schemas are ever involved.

**Proposed fix**

- The conversion depends only on the schema string, so the result can be cached (keyed on that string).

**Measured impact**

- `scan().plan_files()` on an unpartitioned 150-manifest table: **~86 ms → ~48 ms (~1.8× faster)**.
- The saving grows as the number of manifests increases.

I am willing to contribute for this improvement.

贡献指南

这个仓库没有索引到贡献指南

调研方向

从 pyiceberg/avro/file.py 中的 AvroFileHeader.get_schema() 开始,跟踪 avro_to_iceberg 转换。使用重复的 manifest 调用 scan().plan_files(),观察重复的工作,并确认转换结果会按 schema 字符串复用,同时随着 manifest 数量增加,规划性能得到提升。

由索引模型根据 Issue 内容生成。

评估

技术栈
python
领域
performance
Issue 类型
重构
难度
2/5
预计耗时
1-3 小时
活跃度
冷清
描述清晰度
描述清楚
新手友好度
78/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。