Batch ingestion using spark creates duplicates in pinot offline table
Open
- Dominant language
- Java
- Stars
- 6.1k
- Forks
- 1.5k
- Avg merge
- 1d 21h
- Merged PRs (30d)
- 189
Description
This is a query-
trying batch ingestion using spark example spark-3.1.1-bin-hadoop2.7. We see duplicate data in final table. Do we have a workaround for this?
Contributor guide
Research direction
Start with the Spark 3.1.1 batch-ingestion example and reproduce the duplicate data in the Pinot offline table. Gather the ingestion configuration and input details needed to determine whether the duplicates come from the example or the ingestion process; done means documenting a confirmed cause or workaround.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- spark
- Domain
- data-engineering, databases
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100