apache / apache/pinot

Batch ingestion using spark creates duplicates in pinot offline table

Open
#7,002 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
6.1k
Forks
1.5k
Avg merge
1d 21h
Merged PRs (30d)
189

Description

This is a query-

trying batch ingestion using spark example spark-3.1.1-bin-hadoop2.7. We see duplicate data in final table. Do we have a workaround for this?

Contributor guide

Open the contributing guide

Research direction

Start with the Spark 3.1.1 batch-ingestion example and reproduce the duplicate data in the Pinot offline table. Gather the ingestion configuration and input details needed to determine whether the duplicates come from the example or the ingestion process; done means documenting a confirmed cause or workaround.

Written by the indexing model from the issue text.

Assessment

Tech stack
spark
Domain
data-engineering, databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.