apache / apache/hudi

Fix leftover/lingering rollbacks for table services

Open
#16,792 1 comment 0 reactions 1 assignee Claimed by @nsivabalan View on GitHub
area:table-service from-jira priority:high type:bug
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

when a table service fails, and is re-attempted, we trigger a rollback followed by re-attempting the table service. w/ clustering, we have a way to nuke the entire clustering plan. [https://github.com/apache/hudi/blob/baf141abbd6da022c66fa518588e34452a6902b4/hudi-client/hudi-spark-client/src/main/java/org/apache/hudi/table/action/commit/BaseSparkCommitActionExecutor.java#L142] 

 

but there are chances we end up in below state

 

t6.rc.req

t6.rc.inflight 

// trigger rollback. 

t7.rb.req

t7.rb.inflig

delete all data files from t6.rc

write to metadata table 

delete t6.rc* files from timeline 

and we crash. 

 

So, we might have a lingering rollback plan t7.rb.req and t7.rb.inflight in the timeline forever. 

 

This might only be an issue when we try to nuking the entire plan. Otherwise, t6.rc.req will never be cleaned up and next attempt will retry it. 

 

 

 

 

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-8886
- Type: Bug
- Fix version(s):
- 1.1.0

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.