apache / apache/hudi

[SUPPORT] hive-sync

Open
#11,721 5 comments 0 reactions 0 assignees View on GitHub
component:catalog-sync
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

**To Reproduce**

Whether there are parameters in **hive_sync** can be controlled. Each synchronization will only incrementally synchronize the partition contents, and will no longer complete the missing partitions in hive-matestore. Because I will clean up the historical hive partition data to ensure that there is a stable amount of partition data in hive instead of growing all the time.

**Expected behavior**

A clear and concise description of what you expected to happen.

**Environment Description**

* Hudi version : 0.14.1

* Spark version : spark3.3

* Hive version : 3.1.3

* Hadoop version : 3.3.6

* Storage (HDFS/S3/GCS..) : GCS

* Running on Docker? (yes/no) : no

Contributor guide

No contributing guide indexed for this repository

Research direction

The report concerns hive_sync partition synchronization in Hudi 0.14.1 with Spark 3.3, Hive 3.1.3, Hadoop 3.3.6, and GCS, but names no files, tests, or entry points. Start by locating the hive_sync implementation and its partition synchronization tests, then clarify the requested configuration and verify that incremental synchronization does not recreate partitions removed from the metastore.

Written by the indexing model from the issue text.

Assessment

Tech stack
google-cloud, hadoop, java, spark
Domain
data-engineering, databases
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.