4paradigm / 4paradigm/OpenMLDB
optimize the pre-aggregation of out-of-order
- Dominant language
- C++
- Stars
- 1.7k
- Forks
- 331
- Avg merge
- 12d 12h
- Merged PRs (30d)
- 1
Description
**Describe the feature you'd like**
Currently, if the records arriving out of order are not located in the written time interval, they will be written to the table separately, which will affect the performance of pre-aggregation.
**Additional context**
https://github.com/4paradigm/OpenMLDB/blob/687b279e863afceb837d11bbe38bc5c3f37163f6/src/storage/aggregator.cc#L174
Contributor guide
Research direction
The issue points to src/storage/aggregator.cc line 174. Start by understanding the pre-aggregation logic and how out-of-order records are currently handled. Examine the code around that line to see where records outside the written time interval are written separately. Determine what 'optimize' means in this context—likely modifying the aggregation logic to batch or reorder these writes. Run existing tests related to aggregator to ensure changes don't break functionality.
Written by the indexing model from the issue text.
Assessment
- Domain
- databases, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100