4paradigm / 4paradigm/OpenMLDB

load data: load parquet to online, unable to ensure correctness

Offen
#3,424 0 Kommentare 0 Reaktionen 2 zugewiesene Personen Beansprucht von @vagetablechicken Auf GitHub ansehen
batch-engine bug high-priority
Vorherrschende Sprache
C++
Sterne
1.7k
Forks
331
Ø Merge
12 T. 12 Std.
Gemergte PRs (30 T.)
1

Beschreibung

load data can't ensure correctness: insert more rows to online than parquet file

Steps:
1. load data from parquet ( 600 million rows, 3.5 G parquet file in disk, 300 parquet files total )
2. with spark local, driver memory: 64G, max task in parallel: 8
3. spark will split job into 50 tasks, and 8 tasks one time
4. tasks might fail due to `IOException`, rpc to tablet server timeout of 20s
5. spark sees to retry the whole job (50 tasks), result in insert data from parquet double times

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

The issue describes a data loading correctness problem in OpenMLDB involving Spark jobs and Parquet files. Look at the load data flow, likely in the Spark connector or ingestion module, focusing on task failure handling and retry logic. Check for idempotency mechanisms and RPC timeout configurations to tablet servers. Running a test with a smaller dataset to reproduce the duplicate insertion would be a first step.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
spark
Bereich
data-engineering, databases
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.