4paradigm / 4paradigm/OpenMLDB

load data from parquet, poor performance for big datasets

未关闭
#3,396 2 条评论 0 个 reaction 已指派 2 人 已被 @vagetablechicken 认领 在 GitHub 查看
batch-engine bug high-priority
主要语言
C++
星标
1.7k
派生
331
平均合并
12 天 12 小时
30 天内合并 PR
1

描述

parquet file is 3.5G, and possibly 35G in memory. load data from parquet into online storage, spark bootstrap the job, starting 30min and fails of OOM.

贡献指南

打开贡献指南

调研方向

The issue mentions loading a large Parquet file (3.5G, potentially 35G in memory) via Spark into OpenMLDB, with OOM failures after 30 minutes. Start by examining the Spark integration code and memory management for data ingestion. Look for configuration options or limits on chunk size or streaming reads. Determine what 'done' looks like by verifying the data loads without OOM and within a reasonable time.

由索引模型根据 Issue 内容生成。

评估

技术栈
spark
领域
data-engineering, databases, performance
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
需要澄清
新手友好度
30/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。