apache / apache/hudi

Identify out of the box default performance flips for spark-sql

Open
#15,147 0 comments 0 reactions 0 assignees View on GitHub
from-jira priority:high type:improvement
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

We had HUDI-2151 to track performance flips, but its been 1 year that we combed through all configs. Lets do another round of combing through all configs and come up with a new list to flip. 

this ticket specifically tracks spark-sql layer configs. 

 

 

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-4106
- Type: Improvement
- Epic: https://issues.apache.org/jira/browse/HUDI-3249

---

## Comments

15/Sep/22 19:05;alexey.kudinkin;[~shivnarayan] ticket states that there's a patch available, can you please link it here?;;;

Contributor guide

No contributing guide indexed for this repository

Research direction

Review HUDI-2151 and the linked HUDI-4106 context first, then examine the Spark SQL layer's configuration defaults. Compare the current defaults across the configs and produce a new list of performance settings that should be flipped.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
data-engineering, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.