4paradigm / 4paradigm/OpenMLDB

feat: support distributed query on BatchMode under some restrictions

Aperta
#318 0 commenti 0 reazioni 1 assegnatario Rivendicata da @jingchen2222 Vedi su GitHub
enhancement
Lingua principale
C++
Stelle
1.7k
Fork
331
Merge medio
12g 12h
PR unite (30g)
1

Descrizione

**Is your feature request related to a problem? Please describe.**

Users can be frustrated when distributed batch queries aren't suported currently.
We are considering support batch queried under the distributed environment with some specific restrictions.
- simple query can be supported because the queried rows are independent.
- complex query with aggregation can be supported if the aggregation data has been are partitioned by index.
- aggregation query on data group by index
- aggregation query on data filter by index
- aggregation query on window filter by index

In this issue, we are going to support aggregation query on data group/filter by partition key.

**Describe the solution you'd like**

Before we deep into the implementation of distributed batch queries, we have to deal with some related issues below:
- Support general aggeration like, `SUM`, `MIN`, `MAX` etc, on the whole table #219
- Optimize SQL plan when the queried table is filtered and grouped by the same key/expression #317

**Describe alternatives you've considered**
A clear and concise description of any alternative solutions or features you've considered.

**Additional context**
Add any other context or screenshots about the feature request here.

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

The issue describes adding distributed batch query support for aggregation queries filtered/grouped by partition key. Start by reviewing the existing batch query and distributed execution code. Look at related issues #219 and #317 for prerequisite work. Understand the table partitioning and aggregation execution paths. A newcomer would need to understand the distributed query engine and SQL planner.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
sql
Ambito
databases, distributed-systems, machine-learning
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
20/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.