apache / apache/paimon

[Bug] Use the tag incremental query, file does not exist

Open
#1,939 1 comment 0 reactions 0 assignees View on GitHub
bug
Dominant language
Java
Stars
3.4k
Forks
1.4k
Avg merge
1d 11h
Merged PRs (30d)
396

Description

### Search before asking

- [X] I searched in the [issues](https://github.com/apache/incubator-paimon/issues) and found nothing similar.

### Paimon version

0.5

### Compute Engine

flink1.16.1
spark3.3.1

### Minimal reproduce step

schema options
` "options" : {
"owner" : "root",
"partition.expiration-check-interval" : "1 d",
"tag.automatic-creation" : "process-time",
"tag.creation-period" : "daily",
"partition.expiration-time" : "30 d",
"bucket" : "50",
"file.compression" : "ZSTD",
"snapshot.time-retained" : "24 H",
"bucket-key" : "#log_uuid",
"partition.timestamp-formatter" : "yyyy-MM-dd",
"file.format" : "parquet",
"tag.num-retained-max" : "10",
"metadata.stats-mode" : "none",
"tag.creation-delay" : "20 m"
},`

spark sql:
`SELECT count(1) cnt FROM paimon_incremental_query('bdc_ods.ods_log_paimon_inc_1d', '2023-09-02', '2023-09-03');`

### What doesn't meet your expectations?

`Caused by: java.io.FileNotFoundException: File 'oss://bucket_dev/user/hive/warehouse/bdc_ods.db/ods_log_paimon_inc_1d/dt=2023-09-02/bucket-7/data-7d5e1fe1-55a5-4f23-9d0f-57e6d9d63063-2.parquet' not found, Possible causes: 1.snapshot expires too fast, you can configure 'snapshot.time-retained' option with a larger value. 2.consumption is too slow, you can improve the performance of consumption (For example, increasing parallelism).
at org.apache.paimon.utils.FileUtils.createFormatReader(FileUtils.java:119)
at org.apache.paimon.io.KeyValueDataFileRecordReader.(KeyValueDataFileRecordReader.java:55)
at org.apache.paimon.io.KeyValueFileReaderFactory.createRecordReader(KeyValueFileReaderFactory.java:95)
at org.apache.paimon.mergetree.MergeTreeReaders.lambda$readerForRun$2(MergeTreeReaders.java:88)
at org.apache.paimon.mergetree.compact.ConcatRecordReader.create(ConcatRecordReader.java:50)
at org.apache.paimon.mergetree.MergeTreeReaders.readerForRun(MergeTreeReaders.java:91)
at org.apache.paimon.mergetree.MergeTreeReaders.lambda$readerForSection$1(MergeTreeReaders.java:77)
at org.apache.paimon.mergetree.MergeSorter.mergeSort(MergeSorter.java:119)
at org.apache.paimon.mergetree.MergeTreeReaders.readerForSection(MergeTreeReaders.java:79)
at org.apache.paimon.operation.KeyValueFileStoreRead.lambda$batchMergeRead$4(KeyValueFileStoreRead.java:235)
at org.apache.paimon.mergetree.compact.ConcatRecordReader.create(ConcatRecordReader.java:50)
at org.apache.paimon.operation.KeyValueFileStoreRead.batchMergeRead(KeyValueFileStoreRead.java:245)
at org.apache.paimon.operation.KeyValueFileStoreRead.createReaderWithoutOuterProjection(KeyValueFileStoreRead.java:208)
at org.apache.paimon.operation.KeyValueFileStoreRead.createReader(KeyValueFileStoreRead.java:182)
at org.apache.paimon.table.source.KeyValueTableRead.createReader(KeyValueTableRead.java:51)
at org.apache.paimon.spark.SparkReaderFactory.createReader(SparkReaderFactory.java:54)`

### Anything else?

_No response_

### Are you willing to submit a PR?

- [ ] I'm willing to submit a PR!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the Spark SQL incremental query with the provided Paimon options and inspect the stack-trace entry points, especially SparkReaderFactory.java, KeyValueTableRead.java, and KeyValueFileStoreRead.java. Determine why the incremental read references a missing Parquet file; done means the query completes without the reported FileNotFoundException.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering, databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.