4paradigm / 4paradigm/OpenMLDB

abnormal average time when querying different data volume for the same key

未关闭
#3,871 1 条评论 0 个 reaction 已指派 2 人 已被 @aceforeverd 认领 在 GitHub 查看
storage-engine
主要语言
C++
星标
1.7k
派生
331
平均合并
12 天 12 小时
30 天内合并 PR
1

描述

**Description**
During query performance testing, it was found that querying all data rows for a key incurs the least time cost; querying a subset of data within a specified time range for a key results in a relatively increased time cost; querying data for a specific timestamp for a key results in an even greater increase in time cost.

Detail:
table:import TalkingData train dataset (180+ million rows) into openmldb
query key: ip=88

| key | other condition | rows | average time(us) |
|-------|-------------------------------------------------------|-------|------------------|
| ip=88 | all | 4278 | 9774.375 |
| ip=88 | '2017-11-06 00:00:00' <= ts < "2017-11-07 00:00:00" | 183 | 11570.365 |
| ip=88 | ts='2017-11-06 16:19:38' | 1 | 16504.145 |

more result detail: https://qiok3h8ob4.feishu.cn/docx/YkYfdBZm9oVk0MxLFx9co8lLn1g?from=from_copylink

**Expected Behavior**
querying smaller amounts of data should have shorter time costs, or at least not longer than querying larger amounts of data.

**Steps to Reproduce**

1. deploy openmldb;
2. load data (TalkingData train.csv)into table;
3. find a key with large enough total data volume;
4. execute queries and calculate the average time cost;

贡献指南

打开贡献指南

调研方向

The issue describes a performance anomaly in OpenMLDB where querying fewer rows for a key takes longer. Start by examining the query execution path for timestamp-based filters versus full key scans. Look at the indexing and data layout for the 'ts' column. Reproduce the issue using the TalkingData dataset and profile the queries to identify bottlenecks.

由索引模型根据 Issue 内容生成。

评估

技术栈
sql
领域
databases, performance
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。