4paradigm / 4paradigm/OpenMLDB

load data from parquet, poor performance for big datasets

Offen
#3,396 2 Kommentare 0 Reaktionen 2 zugewiesene Personen Beansprucht von @vagetablechicken Auf GitHub ansehen
batch-engine bug high-priority
Vorherrschende Sprache
C++
Sterne
1.7k
Forks
331
Ø Merge
12 T. 12 Std.
Gemergte PRs (30 T.)
1

Beschreibung

parquet file is 3.5G, and possibly 35G in memory. load data from parquet into online storage, spark bootstrap the job, starting 30min and fails of OOM.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

The issue mentions loading a large Parquet file (3.5G, potentially 35G in memory) via Spark into OpenMLDB, with OOM failures after 30 minutes. Start by examining the Spark integration code and memory management for data ingestion. Look for configuration options or limits on chunk size or streaming reads. Determine what 'done' looks like by verifying the data loads without OOM and within a reasonable time.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
spark
Bereich
data-engineering, databases, performance
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
30/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.