apache / apache/hudi

Use table or write schema instead of deducing schema per file group for clustering

Open
#16,441 0 comments 0 reactions 1 assignee Assigned to @yihua View on GitHub
from-jira priority:high type:improvement
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

Right now each clustering group derives the schema on its own. Conceptually we can use one schema for all clustering groups. This is a behavior change and needs revisiting clustering logic end-to-end.

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-7586
- Type: Improvement
- Fix version(s):
- 0.16.0
- 1.1.0

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.