provide an option to skip a page in case corrupted bytes occur
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.6k
- Avg merge
- 3d 12h
- Merged PRs (30d)
- 33
Description
In case of hardware failure (disk, memory, etc), there might be corrupted bytes. That will result in ArrayIndexOutOfBoundException or/and data garbled.
Currently, jobs reading those Parquet files will fail unless the corrupted files are deleted/moved.
Currently page metadata has a CRC field (not used so far), which can be used to check integrity of the page. If page data is corrupted, skip the whole page.
related issue: PARQUET-148
**Reporter**: [Tongjie Chen](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=tongjie) / @tongjiechen
**Note**: *This issue was originally created as [PARQUET-149](https://issues.apache.org/jira/browse/PARQUET-149). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*
Contributor guide
No contributing guide indexed for this repository
Research direction
Start at the Parquet page-reading path and the page metadata handling described in the issue; inspect how the existing CRC field can be checked before page data is consumed. Reproduce corruption or add focused coverage for corrupted page bytes, then verify that the affected page is skipped while the job continues instead of failing or returning garbled data.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100