apache / apache/gravitino

[Subtask] [spark-connector] optimize read hive bucket table

Open
#1,567 0 comments 0 reactions 0 assignees View on GitHub
subtask
Dominant language
Java
Stars
3.2k
Forks
935
Avg merge
1d 17h
Merged PRs (30d)
339

Description

### Describe the subtask

optimize the read with bucket info

### Parent issue

#1227

Contributor guide

Open the contributing guide

Research direction

Start by reading parent issue #1227 and tracing the Spark connector's read path for Hive bucket tables. The issue is complete when bucket information is used to optimize those reads, but it does not name specific files, tests, or acceptance criteria.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.