Issues with AWS EMR + Parquet
Open
- Dominant language
- Java
- Stars
- 6.1k
- Forks
- 1.5k
- Avg merge
- 1d 21h
- Merged PRs (30d)
- 189
Description
Spoke with Dunith Dhanushka and Mayank. They was able to recreate "File does not exist" issues encountered using AWS EMR (spark) to ingest data in parquet format. Mentioned that suspect is the parquet libraries that came with Hadoop dependencies.
Contributor guide
Research direction
Start by reproducing the reported AWS EMR Spark ingestion of Parquet data and examining the Hadoop dependencies involved. Confirm the conditions that produce the "File does not exist" error, then verify that the ingestion completes successfully without that error.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, hadoop, java, spark
- Domain
- data-engineering, distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100