apache / apache/hudi

[SUPPORT]What would be the best or high performant production level hoodie configs to be used for unpartitioned dataset??

Open
#8,820 1 comment 0 reactions 0 assignees View on GitHub
priority:medium type:feature
Dominant language
Java
Stars
6.2k
Forks
2.5k
Avg merge
2d 8h
Merged PRs (30d)
111

Description

**Describe the problem you faced**

A clear and concise description of the problem.
I need to know from hoodie gurus what would be the best configuration for high performing read / write operations. In the given scenario i have multiple files each with average sized of `25MB`. Total size of all files together would be 6GB. Total number of files is 354. Its all JSON data. We want to ingest it into hudi as soon as possible with the hudi metadata.

Please not we don't have fix field for partitioning. So if possible can you tell us configs for un-partitioned data

So the ask is what would be precise hoodie configs which we can use to get quick write and read.

^^ @nsivabalan @vinothchandar @umehrot2 or anybody who may have some insights.

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are named. Start by reviewing Hudi configuration guidance for unpartitioned JSON ingestion and define completion as an agreed production configuration validated against the stated 6GB, 354-file read/write workload.

Written by the indexing model from the issue text.

Assessment

Domain
data-engineering, performance
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.