apache / apache/hudi

Identify out of the box performance config flips for spark-ds

Open
#15,191 0 comments 0 reactions 0 assignees View on GitHub
area:config from-jira priority:critical type:improvement
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

we need to identify out of the box performance flips. Refer to HUDI-2151 for older ticket. But we need to comb through all configs once again and come up with an updated list. 

 

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-4105
- Type: Improvement
- Epic: https://issues.apache.org/jira/browse/HUDI-3249

---

## Comments

15/Sep/22 19:05;alexey.kudinkin;[~shivnarayan] ticket states that there's a patch available, can you please link it here?;;;

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading HUDI-2151 and the linked HUDI-4105, then inspect the Spark data-source configuration options across the project. Compare the current defaults with the older ticket and document an updated list of out-of-the-box performance flips; completion means the list is reviewed and the available patch is linked or its status is clarified.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.