apache / apache/parquet-java

provide an option to skip a page in case corrupted bytes occur

Open
#1,700 0 comments 0 reactions 0 assignees View on GitHub
Component: Parquet Priority: Major Type: enhancement
Dominant language
Java
Stars
3.1k
Forks
1.6k
Avg merge
3d 12h
Merged PRs (30d)
33

Description

In case of hardware failure (disk, memory, etc), there might be corrupted bytes. That will result in ArrayIndexOutOfBoundException or/and data garbled.

Currently, jobs reading those Parquet files will fail unless the corrupted files are deleted/moved.

Currently page metadata has a CRC field (not used so far), which can be used to check integrity of the page. If page data is corrupted, skip the whole page.

related issue: PARQUET-148

**Reporter**: [Tongjie Chen](https://issues.apache.org/jira/secure/ViewProfile.jspa?name=tongjie) / @tongjiechen

**Note**: *This issue was originally created as [PARQUET-149](https://issues.apache.org/jira/browse/PARQUET-149). Please see the [migration documentation](https://issues.apache.org/jira/browse/PARQUET-2502) for further details.*

Contributor guide

No contributing guide indexed for this repository

Research direction

Start at the Parquet page-reading path and the page metadata handling described in the issue; inspect how the existing CRC field can be checked before page data is consumed. Reproduce corruption or add focused coverage for corrupted page bytes, then verify that the affected page is skipped while the job continues instead of failing or returning garbled data.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.