[SUPPORT] hive-sync
- Dominant language
- Java
- Stars
- 6.2k
- Forks
- 2.5k
- Avg merge
- 2d 8h
- Merged PRs (30d)
- 111
Description
**To Reproduce**
Whether there are parameters in **hive_sync** can be controlled. Each synchronization will only incrementally synchronize the partition contents, and will no longer complete the missing partitions in hive-matestore. Because I will clean up the historical hive partition data to ensure that there is a stable amount of partition data in hive instead of growing all the time.
**Expected behavior**
A clear and concise description of what you expected to happen.
**Environment Description**
* Hudi version : 0.14.1
* Spark version : spark3.3
* Hive version : 3.1.3
* Hadoop version : 3.3.6
* Storage (HDFS/S3/GCS..) : GCS
* Running on Docker? (yes/no) : no
Contributor guide
No contributing guide indexed for this repository
Research direction
The report concerns hive_sync partition synchronization in Hudi 0.14.1 with Spark 3.3, Hive 3.1.3, Hadoop 3.3.6, and GCS, but names no files, tests, or entry points. Start by locating the hive_sync implementation and its partition synchronization tests, then clarify the requested configuration and verify that incremental synchronization does not recreate partitions removed from the metastore.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- google-cloud, hadoop, java, spark
- Domain
- data-engineering, databases
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100