4paradigm / 4paradigm/OpenMLDB

feat: support distributed query on BatchMode under some restrictions

Đang mở
#318 0 bình luận 0 reaction 1 người được giao Được @jingchen2222 nhận Xem trên GitHub
enhancement
Ngôn ngữ chính
C++
Star
1.7k
Fork
331
Merge trung bình
12 ngày 12 giờ
Pull request đã merge (30 ngày)
1

Mô tả

**Is your feature request related to a problem? Please describe.**

Users can be frustrated when distributed batch queries aren't suported currently.
We are considering support batch queried under the distributed environment with some specific restrictions.
- simple query can be supported because the queried rows are independent.
- complex query with aggregation can be supported if the aggregation data has been are partitioned by index.
- aggregation query on data group by index
- aggregation query on data filter by index
- aggregation query on window filter by index

In this issue, we are going to support aggregation query on data group/filter by partition key.

**Describe the solution you'd like**

Before we deep into the implementation of distributed batch queries, we have to deal with some related issues below:
- Support general aggeration like, `SUM`, `MIN`, `MAX` etc, on the whole table #219
- Optimize SQL plan when the queried table is filtered and grouped by the same key/expression #317

**Describe alternatives you've considered**
A clear and concise description of any alternative solutions or features you've considered.

**Additional context**
Add any other context or screenshots about the feature request here.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

The issue describes adding distributed batch query support for aggregation queries filtered/grouped by partition key. Start by reviewing the existing batch query and distributed execution code. Look at related issues #219 and #317 for prerequisite work. Understand the table partitioning and aggregation execution paths. A newcomer would need to understand the distributed query engine and SQL planner.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
sql
Lĩnh vực
databases, distributed-systems, machine-learning
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
20/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.