apache / apache/paimon

[Bug] [flink] Could not read External Tables

Open
#1,908 0 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Java
Stars
3.4k
Forks
1.4k
Avg merge
1d 11h
Merged PRs (30d)
396

Description

### Search before asking

- [X] I searched in the [issues](https://github.com/apache/incubator-paimon/issues) and found nothing similar.

### Paimon version

0.5

### Compute Engine

Flink 1.16.2

### Minimal reproduce step

[upload_excel.xlsx](https://github.com/apache/incubator-paimon/files/12459344/upload_excel.xlsx)
1. I trans this excel to parquet

df = pd.read_excel(args.file, sheet_name='data')
parquet_file = '/tmp/parquet_file/' + args.name + '.parquet'
df.to_parquet(parquet_file, index=False)`

2. I upload this parquet to "hdfs://draco01:9870/tmp/flink1.17.1/upload_parquet/table_name.parquet"
3. I create external table :

CREATE TABLE default_database.table_name (
col1 string,
col2 bigint,
col3 string,
col4 date,
col5 string,
PRIMARY KEY (col1,col2) NOT ENFORCED
) WITH (
'connector' = 'paimon',
'path' = 'hdfs://draco01:9870/tmp/flink1.17.1/upload_parquet/table_name.parquet'
);

4. Could read the schema:
+------+--------+-------+-----------------+--------+-----------+
| name | type | null | key | extras | watermark |
+------+--------+-------+-----------------+--------+-----------+
| col1 | STRING | FALSE | PRI(col1, col2) | | |
| col2 | BIGINT | FALSE | PRI(col1, col2) | | |
| col3 | STRING | TRUE | | | |
| col4 | DATE | TRUE | | | |
| col5 | STRING | TRUE | | | |
+------+--------+-------+-----------------+--------+-----------+
5 rows in set
5. Could not read the data, it shows below:
`Caused by: org.apache.hadoop.ipc.RpcException: RPC response exceeds maximum data length`

### What doesn't meet your expectations?

Want to know how to fix it.

### Anything else?

_No response_

### Are you willing to submit a PR?

- [ ] I'm willing to submit a PR!

Contributor guide

No contributing guide indexed for this repository

Research direction

The report provides upload_excel.xlsx, a pandas-to-Parquet conversion, an HDFS path, and the Flink SQL CREATE TABLE as reproduction entry points. Reproduce the read with Paimon 0.5 and Flink 1.16.2, then trace the “RPC response exceeds maximum data length” error; done when the external table data can be read without it.

Written by the indexing model from the issue text.

Assessment

Tech stack
hadoop, java, python
Domain
data-engineering, databases, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.