apache / apache/paimon

[Bug] [flink] Could not read External Tables

未关闭
#1,908 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
bug
主要语言
Java
星标
3.4k
派生
1.4k
平均合并
1 天 9 小时
30 天内合并 PR
423

描述

### Search before asking

- [X] I searched in the [issues](https://github.com/apache/incubator-paimon/issues) and found nothing similar.

### Paimon version

0.5

### Compute Engine

Flink 1.16.2

### Minimal reproduce step

[upload_excel.xlsx](https://github.com/apache/incubator-paimon/files/12459344/upload_excel.xlsx)
1. I trans this excel to parquet

df = pd.read_excel(args.file, sheet_name='data')
parquet_file = '/tmp/parquet_file/' + args.name + '.parquet'
df.to_parquet(parquet_file, index=False)`

2. I upload this parquet to "hdfs://draco01:9870/tmp/flink1.17.1/upload_parquet/table_name.parquet"
3. I create external table :

CREATE TABLE default_database.table_name (
col1 string,
col2 bigint,
col3 string,
col4 date,
col5 string,
PRIMARY KEY (col1,col2) NOT ENFORCED
) WITH (
'connector' = 'paimon',
'path' = 'hdfs://draco01:9870/tmp/flink1.17.1/upload_parquet/table_name.parquet'
);

4. Could read the schema:
+------+--------+-------+-----------------+--------+-----------+
| name | type | null | key | extras | watermark |
+------+--------+-------+-----------------+--------+-----------+
| col1 | STRING | FALSE | PRI(col1, col2) | | |
| col2 | BIGINT | FALSE | PRI(col1, col2) | | |
| col3 | STRING | TRUE | | | |
| col4 | DATE | TRUE | | | |
| col5 | STRING | TRUE | | | |
+------+--------+-------+-----------------+--------+-----------+
5 rows in set
5. Could not read the data, it shows below:
`Caused by: org.apache.hadoop.ipc.RpcException: RPC response exceeds maximum data length`

### What doesn't meet your expectations?

Want to know how to fix it.

### Anything else?

_No response_

### Are you willing to submit a PR?

- [ ] I'm willing to submit a PR!

贡献指南

这个仓库没有索引到贡献指南

调研方向

报告提供了 upload_excel.xlsx、pandas-to-Parquet 转换、HDFS 路径以及 Flink SQL CREATE TABLE,作为复现入口。使用 Paimon 0.5 和 Flink 1.16.2 复现读取过程,然后跟踪“RPC response exceeds maximum data length”错误;当无需该错误即可读取外部表数据时,即完成。

由索引模型根据 Issue 内容生成。

评估

技术栈
hadoop, java, python
领域
data-engineering, databases, distributed-systems
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
30/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。