Parquet metadata check limit optimization
- Dominant language
- Scala
- Stars
- 1.6k
- Forks
- 657
- Avg merge
- 2d 14h
- Merged PRs (30d)
- 80
Description
### Description
Now the validation metadata config(spark.gluten.sql.fallbackUnexpectedMetadataParquet) is default false, if set to true, for each root path, we check the file limit (spark.gluten.sql.fallbackUnexpectedMetadataParquet.limit), if the number of partitions are too much, the validation will be expensive.
The possible solution is to sample the rootPaths to select some files.
The sample file limit should be decided by file total limit number and the total file number in root paths, the latter should be decided by the percentage.
### Gluten version
None
Contributor guide
Research direction
Start by tracing the validation controlled by spark.gluten.sql.fallbackUnexpectedMetadataParquet and its limit setting. Determine how root paths, partition counts, and file totals are currently evaluated; done should mean sampling avoids expensive validation while respecting both the configured total file limit and the requested percentage.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- scala
- Domain
- data
- Issue type
- Refactor
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100