apache / apache/hudi

Enable partial updates for CDC work payload

Open
#16,354 2 comments 0 reactions 0 assignees View on GitHub
from-jira priority:high type:devtask
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

OLTP workloads on upstream databases, often update/delete/insert different columns in the table on each operation. Currently, Hudi can only supporting partial updates in cases where the same columns are being mutated in a given write to Hudi (e.g Spark SQL ETLs with MIT or Update statements). Here, we explore what it takes to support a smarter storage format, that can only encode the changed columns into log along with the different implementations.
h2. Goals
# Enable partial update functionality for all existing and potential future CDC workloads without huge modification or duplication.
# Performance parity with current full-record updates or partial updates across the same set of columns
# Exhibit reduction in storage costs, by only storing the changed columns.
# Should also result in computation cost reductions by scanning/processing less data
# Should not affect the scalability of the existing system ingestion system. The number of files generated for partial update should not increase dramatically.

 

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-7229
- Type: Task
- Epic: https://issues.apache.org/jira/browse/HUDI-6242
- Fix version(s):
- 1.2.0

---

## Comments

03/May/24 02:13;vinoth;Punting this to 1.1 

# [1.1] Implement support on top of data blocks.
## we need to pass change columns information and operation all the way to write handles, using a field in HoodieRecord
## ... 
# [1.1] Implement support on top of cdc data blocks.
## we can track similar bitmaps for cdc data blocks as well
## we need to extend the new file group reader to also merge base and cdc blocks. (not just base and data blocks).;;;

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by tracing how HoodieRecord carries operation and changed-column information through the write handles, then inspect the new file group reader and the existing base, data, and CDC data-block paths. Done means CDC workloads can encode and merge only changed columns, while meeting the stated performance, storage, computation, and file-generation goals.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering, databases, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.