apache / apache/hudi

Improve Documentation for spark-sql

Open
#15,538 0 comments 0 reactions 0 assignees View on GitHub
area:dev-experience area:sql docs from-jira priority:high type:improvement
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

The documentation does not show examples of some use cases, especially populating tables with large amounts of data. Additionally, it is unclear which configs work in spark-sql and which configs do nothing

## JIRA info

- Link: https://issues.apache.org/jira/browse/HUDI-5267
- Type: Improvement
- Fix version(s):
- 0.16.0

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the spark-sql documentation and the linked HUDI-5267 issue. Identify the existing guidance for populating tables and the configuration references, then document large-data examples and clarify which configurations apply to spark-sql; done means both use cases and supported configurations are clearly covered.

Written by the indexing model from the issue text.

Assessment

Tech stack
spark, sql
Domain
data-engineering, documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.