4paradigm / 4paradigm/OpenMLDB

`load data` in cluster mode may cause out-of-order insertion, leading poor performance for long window

Open
#2,544 1 comment 0 reactions 1 assignee Claimed by @zhanghaohit View on GitHub
bug
Dominant language
C++
Stars
1.7k
Forks
331
Avg merge
12d 12h
Merged PRs (30d)
1

Description

**Bug Description**
if `spark` is configured with multiple threads, say `local[*]` or `local[32]`, data loading will run in parallel, causing out-of-order loading.

It is ok for normal data insertion. But `long window optimization` requires that data is loaded mostly in order.

**Potential solution**
- remove the `in order loading` requirement for `long window`, by deleting the old entry and inserting the new one in the pre-aggr table.

Contributor guide

Open the contributing guide

Research direction

The issue relates to the 'load data' operation in cluster mode and its interaction with long window optimization. Investigate the data loading path in the codebase, particularly looking for parallel execution logic. Examine the long window optimization implementation to understand the in-order requirement. Check if there are existing tests for data loading or window performance to verify changes.

Written by the indexing model from the issue text.

Assessment

Tech stack
spark
Domain
databases, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.