apache / apache/pinot

Issues with AWS EMR + Parquet

Open
#7,986 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
6.1k
Forks
1.5k
Avg merge
1d 21h
Merged PRs (30d)
189

Description

Spoke with Dunith Dhanushka and Mayank. They was able to recreate "File does not exist" issues encountered using AWS EMR (spark) to ingest data in parquet format. Mentioned that suspect is the parquet libraries that came with Hadoop dependencies.

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the reported AWS EMR Spark ingestion of Parquet data and examining the Hadoop dependencies involved. Confirm the conditions that produce the "File does not exist" error, then verify that the ingestion completes successfully without that error.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, hadoop, java, spark
Domain
data-engineering, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.