apache / apache/kyuubi

[Bug] DBeaver is slow to obtain and display data

Open
#6,796 2 comments 0 reactions 0 assignees View on GitHub
kind:bug priority:major
Dominant language
Scala
Stars
2.4k
Forks
1k
PR merge metrics
No merged PRs in 30d

Description

### Code of Conduct

- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

### Search before asking

- [X] I have searched in the [issues](https://github.com/apache/kyuubi/issues?q=is%3Aissue) and found no similar issues.

### Describe the bug

The default query of dbeaver displays 200 rows. Each pull-down will display 200 more rows of data. There is a table with a total of 990 data.
1. Every time I pull down to display more data, it will generate new SQL to query. Is this as expected? Is it possible to get the results of the previous query?
2. Each time you pull down to display more data, the time it takes to return and display the data will increase a lot.

image

image

### Affects Version(s)

1.10.0

### Kyuubi Server Log Output

_No response_

### Kyuubi Engine Log Output

_No response_

### Kyuubi Server Configurations

_No response_

### Kyuubi Engine Configurations

```yaml
spark.master yarn
spark.yarn.queue default
spark.executor.cores 1
spark.driver.memory 3g
spark.executor.memory 3g
spark.dynamicAllocation.enabled true
spark.dynamicAllocation.shuffleTracking.enabled true
spark.dynamicAllocation.minExecutors 1
spark.dynamicAllocation.maxExecutors 10
spark.dynamicAllocation.initialExecutors 1
spark.cleaner.periodicGC.interval 5min
```

### Additional context

_No response_

### Are you willing to submit PR?

- [ ] Yes. I would be willing to submit a PR with guidance from the Kyuubi community to fix.
- [ ] No. I cannot submit a PR at this time.

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the DBeaver scrolling behavior against Kyuubi 1.10.0 with the supplied Spark-on-YARN configuration. Capture the SQL generated for each 200-row fetch and compare response and display times as the result window grows. Done means the repeated-query behavior and slowdown have a confirmed cause and a tested resolution or documented limitation.

Written by the indexing model from the issue text.

Assessment

Tech stack
scala, spark
Domain
backend, databases, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.