apache / apache/kyuubi

[Umbrella] Improvements and evaluation for TRowSet generation of Spark Engine

Open
#5,808 0 comments 1 reaction 0 assignees View on GitHub
kind:umbrella priority:major
Dominant language
Scala
Stars
2.4k
Forks
1k
PR merge metrics
No merged PRs in 30d

Description

### Code of Conduct

- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

### Search before asking

- [X] I have searched in the [issues](https://github.com/apache/kyuubi/issues) and found no similar issues.

### Describe the proposal

RowSet generation is that taking the results from result iterator and serializing them into column-based or row-based TRowSet, which is the key point for transportation and performance in most common cases.

1. It's been reported possibility drawbacks in looping the result iterator by wrapped stream in `SparkOperation` (https://github.com/apache/kyuubi/blob/master/externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/operation/SparkOperation.scala#L253C21-L253C26)
2. the performance of Spark Engine's RowSet.toTRowSet should be evaluated by benchmarks, for overall performance and for each data type with different mode of Thrift version and arrow based.
3. Performance Improvements in Spark Engine's RowSet implementation
4. Code cleanup in Spark Engine's RowSet generation

### Task list

- [ ] benchmark ut
- #5809
- [ ] Add benchmark unit test for RowSet generation covering supported data types
- [ ] Add benchmark dedicated unit test for each supported data type for RowSet generation
- [ ] Replace looping the iterator from `toSeq` (`toStream` of Iterator) to immutable collection
- [ ] Compare toStream/toSeq/toList/toVector
- ~~#5804~~
- [ ] Parallel processing for column-based TRowSet generation
- [ ] Performance improvements in data types
- #5811
- ~~DecimalType with column-based mode #5810~~
- ~~ArrayType of primitive data types with column-based mode~~
- Generalize TRowSet generator
- #5851
- #5861
- [ ] Code cleanup in RowSet of Spark Engine
- #5831

### Are you willing to submit PR?

- [X] Yes. I would be willing to submit a PR with guidance from the Kyuubi community to improve.
- [ ] No. I cannot submit a PR at this time.

Contributor guide

Open the contributing guide

Research direction

Start with SparkOperation.scala around the wrapped result iterator at the linked line, then review the related tasks #5809, #5811, #5851, #5861, and #5831. Run or add the proposed RowSet benchmarks for supported data types and compare the listed collection strategies and Thrift or Arrow modes. Done means a specific improvement is benchmarked, implemented, and covered by the relevant task.

Written by the indexing model from the issue text.

Assessment

Tech stack
scala, spark
Domain
backend, data-engineering
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.