Implement Schema on Read for new spark integration
Open
engine:spark
from-jira
priority:high
type:improvement
- Dominant language
- Java
- Stars
- 6.2k
- Forks
- 2.5k
- Avg merge
- 2d 8h
- Merged PRs (30d)
- 111
Description
schema on read is not currently supported for the mor bootstrap file format
## JIRA info
- Link: https://issues.apache.org/jira/browse/HUDI-6611
- Type: Improvement
- Epic: https://issues.apache.org/jira/browse/HUDI-6568
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading the linked JIRA issue and the existing new Spark integration related to the MOR bootstrap file format. Trace how schemas are handled during bootstrap and identify the expected schema-on-read behavior. Done means schema-on-read is supported for this format and the relevant integration behavior is covered by tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, spark
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100