apache / apache/kyuubi

[Bug] pyspark code will get stuck in beeline

Open
#5,142 8 comments 0 reactions 0 assignees View on GitHub
kind:bug priority:major
Dominant language
Scala
Stars
2.4k
Forks
1k
PR merge metrics
No merged PRs in 30d

Description

### Code of Conduct

- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

### Search before asking

- [X] I have searched in the [issues](https://github.com/apache/kyuubi/issues?q=is%3Aissue) and found no similar issues.

### Describe the bug

simple example:

```shell
SET kyuubi.operation.language=python;
import pandas as pd
from pyspark.sql.functions import pandas_udf
spark.createDataFrame(pd.DataFrame([1, 2, 3], columns=["v"])).collect();
```

The pyspark task will get stuck.
image

thread dump in executor:
image

### Affects Version(s)

master

### Kyuubi Server Log Output

_No response_

### Kyuubi Engine Log Output

_No response_

### Kyuubi Server Configurations

_No response_

### Kyuubi Engine Configurations

_No response_

### Additional context

_No response_

### Are you willing to submit PR?

- [ ] Yes. I would be willing to submit a PR with guidance from the Kyuubi community to fix.
- [ ] No. I cannot submit a PR at this time.

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the supplied beeline command sequence with the Python and Spark example, then inspect the executor thread dump shown in the issue. Done means the PySpark task completes instead of remaining stuck; the issue does not name a source file or test to modify.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, spark
Domain
backend, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.