[Subtask] [spark-connector] support parquet&orc vector read in hive
Open
subtask
- Dominant language
- Java
- Stars
- 3.2k
- Forks
- 935
- Avg merge
- 1d 15h
- Merged PRs (30d)
- 315
Description
### Describe the subtask
use parquet&orc vector read to improment the performance.
### Parent issue
#1227
Contributor guide
Research direction
Start with the parent issue #1227 and the Spark connector's Hive integration; the payload does not name specific files, tests, or entry points. Determine how Parquet and ORC vector reads are expected to work and how performance should be measured. Done means vector reads are supported for both formats with improved performance.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java, spark
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100