apache / apache/hudi

Issue during read from Redshift (via Spectrum) over a Hudi Table - AWS

Open
#18,272 0 comments 0 reactions 0 assignees View on GitHub
type:bug
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

### Bug Description

What happened:
We are facing a problem in a Data Platform based over AWS Cloud Infrastructure. After upgrading the EMR cluster from version 6.9 to 7.10, we are experiencing an issue with CDC-updated tables stored on S3. When reading the data through Redshift procedures, some records are not returned even though their Hudi commit time is earlier than the query execution time.

What you expected:
Some records with a Hudi commit time earlier than the query time should be correctly read and returned by Redshift (probably data under the same partition).

Steps to reproduce:

1. Upgrade the EMR cluster from version 6.9 to 7.10.
2. Run a PySpark script that implements CDC logic and writes data to S3 (Hudi tables).
3. While the PySpark CDC job is running (or after it completes), query the same data using Redshift procedures and observe that some records are not returned despite having a commit time earlier than the query time.

type of write operation in pyspark:
df_upsert_u_clean.write.format('org.apache.hudi') \
.option('hoodie.datasource.write.operation', 'upsert') \
.options(**hudiOptions_dfkkop) \
.mode('append') \
.save(S3_OUTPUT_PATH)

### Environment

**Hudi version:** 0.15.0-amzn-7
**Query engine:** PySpark
**Relevant configs:**
'hoodie.table.name': TABLENAME,
'hoodie.datasource.write.table.type': 'COPY_ON_WRITE',
'hoodie.datasource.write.recordkey.field': ','.join(key_cols),
'hoodie.datasource.write.keygenerator.class': 'org.apache.hudi.keygen.ComplexKeyGenerator',
'hoodie.datasource.write.partitionpath.field': 'PATH_KEY',
'hoodie.datasource.write.precombine.field': 'TSTMP',
'hoodie.index.type':'BLOOM',
'hoodie.simple.index.update.partition.path':'true',
'hoodie.datasource.write.hive_style_partitioning':'true',
'hoodie.datasource.hive_sync.enable': 'true',
'hoodie.datasource.hive_sync.table': TABLENAME,
'hoodie.datasource.hive_sync.database': 'ods',
'hoodie.datasource.hive_sync.partition_fields': 'PATH_KEY',
'hoodie.datasource.hive_sync.partition_extractor_class': 'org.apache.hudi.hive.MultiPartKeysValueExtractor',
'hoodie.datasource.hive_sync.use_jdbc': 'false',
'hoodie.datasource.hive_sync.mode': 'hms',
'hoodie.copyonwrite.record.size.estimate':3761,
'hoodie.parquet.small.file.limit': 104857600,
'hoodie.parquet.max.file.size': 120000000,
'hoodie.datasource.write.schema.evolution.enable': 'true'

### Logs and Stack Trace

_No response_

Contributor guide

No contributing guide indexed for this repository

Research direction

No repository files or tests are named. Start by reproducing the CDC upsert and Redshift read across EMR 6.9 and 7.10, then inspect Hudi commit and partition visibility for the missing records. Done requires identifying the compatibility or query-read cause and defining a verified fix or regression test.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, python, spark
Domain
cloud, data-engineering, databases, distributed-systems
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.