apache / apache/seatunnel

[Feature][connector-clickhouse-v2] Support Split Clickhouse Sql in DynamicChunkSpliiter

Open
#9,672 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
9.7k
Forks
2.4k
Avg merge
3d 9h
Merged PRs (30d)
204

Description

### Search before asking

- [x] I had searched in the [feature](https://github.com/apache/seatunnel/issues?q=is%3Aissue+label%3A%22Feature%22) and found no similar feature requirement.

### Description

In the previous pr (#9446), concurrent reads of clickhouse based on both `table_path` and `query_sql` were implemented. For the `query_sql` mode, the current implementation only concurrently executes the sql in each local table, and the concurrency depends on the number of cluster nodes.

Is it still necessary to implement the DynamicChunkSplitter mode under the `query_sql` mode to split the sql for concurrent queries? If necessary, I will refer to the implementation of JDBC's DynamicChunkSpliiter to achieve dynamic splitting for Clickhouse Sql.

### Usage Scenario

_No response_

### Related issues

_No response_

### Are you willing to submit a PR?

- [x] Yes I am willing to submit a PR!

### Code of Conduct

- [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing previous PR #9446 and the JDBC DynamicChunkSplitter implementation referenced in the issue. Clarify whether query_sql mode should support DynamicChunkSplitter for ClickHouse, then define completion as dynamically splitting ClickHouse SQL for concurrent queries.

Written by the indexing model from the issue text.

Assessment

Tech stack
clickhouse, java
Domain
data-engineering, databases
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.