4paradigm / 4paradigm/OpenMLDB
feat: support distributed query on BatchMode under some restrictions
- 主要语言
- C++
- 星标
- 1.7k
- 派生
- 331
- 平均合并
- 12 天 12 小时
- 30 天内合并 PR
- 1
描述
**Is your feature request related to a problem? Please describe.**
Users can be frustrated when distributed batch queries aren't suported currently.
We are considering support batch queried under the distributed environment with some specific restrictions.
- simple query can be supported because the queried rows are independent.
- complex query with aggregation can be supported if the aggregation data has been are partitioned by index.
- aggregation query on data group by index
- aggregation query on data filter by index
- aggregation query on window filter by index
In this issue, we are going to support aggregation query on data group/filter by partition key.
**Describe the solution you'd like**
Before we deep into the implementation of distributed batch queries, we have to deal with some related issues below:
- Support general aggeration like, `SUM`, `MIN`, `MAX` etc, on the whole table #219
- Optimize SQL plan when the queried table is filtered and grouped by the same key/expression #317
**Describe alternatives you've considered**
A clear and concise description of any alternative solutions or features you've considered.
**Additional context**
Add any other context or screenshots about the feature request here.
贡献指南
调研方向
The issue describes adding distributed batch query support for aggregation queries filtered/grouped by partition key. Start by reviewing the existing batch query and distributed execution code. Look at related issues #219 and #317 for prerequisite work. Understand the table partitioning and aggregation execution paths. A newcomer would need to understand the distributed query engine and SQL planner.
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- sql
- 领域
- databases, distributed-systems, machine-learning
- Issue 类型
- 功能
- 难度
- 5/5
- 预计耗时
- 一周以上
- 活跃度
- 停滞
- 描述清晰度
- 基本清楚
- 新手友好度
- 20/100