apache / apache/hudi

Spark: Support PruneColumn for DataSourceV2 Read

Open
#15,346 0 comments 0 reactions 0 assignees View on GitHub
from-jira priority:high type:improvement
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 4h
Merged PRs (30d)
112

Description

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-4641
- Type: Improvement
- Epic: https://issues.apache.org/jira/browse/HUDI-4449

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the linked HUDI-4641 JIRA issue and the project's Spark DataSourceV2 read entry points; no source files or tests are named here. Determine where column pruning is handled and define completion as Spark reads supporting PruneColumn for DataSourceV2.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.