4paradigm / 4paradigm/OpenMLDB

abnormal average time when querying different data volume for the same key

Offen
#3,871 1 Kommentar 0 Reaktionen 2 zugewiesene Personen Beansprucht von @aceforeverd Auf GitHub ansehen
storage-engine
Vorherrschende Sprache
C++
Sterne
1.7k
Forks
331
Ø Merge
12 T. 12 Std.
Gemergte PRs (30 T.)
1

Beschreibung

**Description**
During query performance testing, it was found that querying all data rows for a key incurs the least time cost; querying a subset of data within a specified time range for a key results in a relatively increased time cost; querying data for a specific timestamp for a key results in an even greater increase in time cost.

Detail:
table:import TalkingData train dataset (180+ million rows) into openmldb
query key: ip=88

| key | other condition | rows | average time(us) |
|-------|-------------------------------------------------------|-------|------------------|
| ip=88 | all | 4278 | 9774.375 |
| ip=88 | '2017-11-06 00:00:00' <= ts < "2017-11-07 00:00:00" | 183 | 11570.365 |
| ip=88 | ts='2017-11-06 16:19:38' | 1 | 16504.145 |

more result detail: https://qiok3h8ob4.feishu.cn/docx/YkYfdBZm9oVk0MxLFx9co8lLn1g?from=from_copylink

**Expected Behavior**
querying smaller amounts of data should have shorter time costs, or at least not longer than querying larger amounts of data.

**Steps to Reproduce**

1. deploy openmldb;
2. load data (TalkingData train.csv)into table;
3. find a key with large enough total data volume;
4. execute queries and calculate the average time cost;

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

The issue describes a performance anomaly in OpenMLDB where querying fewer rows for a key takes longer. Start by examining the query execution path for timestamp-based filters versus full key scans. Look at the indexing and data layout for the 'ts' column. Reproduce the issue using the TalkingData dataset and profile the queries to identify bottlenecks.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
sql
Bereich
databases, performance
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.