apache / apache/gluten

[VL] Gluten runs slower for some TPC-H queries than Vanilla Spark

Open
#8,466 4 comments 0 reactions 0 assignees View on GitHub
bug triage
Dominant language
Scala
Stars
1.6k
Forks
657
Avg merge
2d 14h
Merged PRs (30d)
80

Description

### Backend

VL (Velox)

### Bug description

I am testing Gluten with Velox backend for the given TPC-H benchmark scripts provided in the repo. It is observed that few SQL queries q7, q9, q10, q12 runs slower with gluten.
What is the reason for the slower performance for these queries and how to improve them?
I am running the tests on ARM based AWS instance :
m7g.4xlarge , VCPUs = 16, Memory = 64GB
Spark Version : 3.5.2
Data size : Used scale factor SF=100
Below is the shell script used to run the tests:
**For Gluten**

```
GLUTEN_JAR=/path/to/incubator-gluten/package/target/gluten-velox-bundle-spark3.5_2.12-ubuntu_22.04_aarch_64-1.3.0-SNAPSHOT.jar
SPARK_HOME=/home/spark/spark-3.5.2
cat tpch_parquet.scala | ${SPARK_HOME}/bin/spark-shell \
--master spark://172.32.5.244:7077 --deploy-mode client \
--conf spark.plugins=org.apache.gluten.GlutenPlugin \
--conf spark.driver.extraClassPath=${GLUTEN_JAR} \
--conf spark.executor.extraClassPath=${GLUTEN_JAR} \
--conf spark.memory.offHeap.enabled=true \
--conf spark.memory.offHeap.size=12g \
--conf spark.gluten.sql.columnar.forceShuffledHashJoin=true \
--conf spark.driver.memory=4G \
--conf spark.executor.instances=1 \
--conf spark.executor.memory=30G \
--conf spark.executor.cores=16 \
--conf spark.executor.memoryOverhead=2g \
--conf spark.driver.maxResultSize=2g \
--conf spark.shuffle.manager=org.apache.spark.shuffle.sort.ColumnarShuffleManager \
--conf spark.driver.extraJavaOptions="--illegal-access=permit -Dio.netty.tryReflectionSetAccessible=true --add-opens java.base/java.lang=ALL-UNNAMED --add-opens java.base/java.util=ALL-UNNAMED" \
--conf spark.executor.extraJavaOptions="--illegal-access=permit -Dio.netty.tryReflectionSetAccessible=true --add-opens java.base/java.lang=ALL-UNNAMED --add-opens java.base/java.util=ALL-UNNAMED" \
```

**For Vanilla Spark**
```
SPARK_HOME=/home/spark/spark-3.5.2
cat tpch_parquet.scala | ${SPARK_HOME}/bin/spark-shell \
--master spark://172.32.5.244:7077 --deploy-mode client \
--conf spark.memory.offHeap.enabled=true \
--conf spark.memory.offHeap.size=12g \
--conf spark.driver.memory=4G \
--conf spark.executor.instances=1 \
--conf spark.executor.memory=30G \
--conf spark.executor.cores=16 \
--conf spark.executor.memoryOverhead=2g \
--conf spark.driver.maxResultSize=2g \
```

![image](https://github.com/user-attachments/assets/c9f2e585-2fb0-4215-bde0-7562b4d4ab5a)

### Spark version

Spark-3.5.x

### Spark configurations

GLUTEN_JAR=/path/to/incubator-gluten/package/target/gluten-velox-bundle-spark3.5_2.12-ubuntu_22.04_aarch_64-1.3.0-SNAPSHOT.jar
SPARK_HOME=/home/spark/spark-3.5.2
cat tpch_parquet.scala | ${SPARK_HOME}/bin/spark-shell \
--master spark://172.32.5.244:7077 --deploy-mode client \
--conf spark.plugins=org.apache.gluten.GlutenPlugin \
--conf spark.driver.extraClassPath=${GLUTEN_JAR} \
--conf spark.executor.extraClassPath=${GLUTEN_JAR} \
--conf spark.memory.offHeap.enabled=true \
--conf spark.memory.offHeap.size=12g \
--conf spark.gluten.sql.columnar.forceShuffledHashJoin=true \
--conf spark.driver.memory=4G \
--conf spark.executor.instances=1 \
--conf spark.executor.memory=30G \
--conf spark.executor.cores=16 \
--conf spark.executor.memoryOverhead=2g \
--conf spark.driver.maxResultSize=2g \
--conf spark.shuffle.manager=org.apache.spark.shuffle.sort.ColumnarShuffleManager \
--conf spark.driver.extraJavaOptions="--illegal-access=permit -Dio.netty.tryReflectionSetAccessible=true --add-opens java.base/java.lang=ALL-UNNAMED --add-opens java.base/java.util=ALL-UNNAMED" \
--conf spark.executor.extraJavaOptions="--illegal-access=permit -Dio.netty.tryReflectionSetAccessible=true --add-opens java.base/java.lang=ALL-UNNAMED --add-opens java.base/java.util=ALL-UNNAMED" \

### System information

Gluten Version: 1.3.0-SNAPSHOT
Commit: 4dfdfd77b52f2f98fa0cf32eca143b47e4bd11b5
CMake Version: 3.28.3
System: Linux-6.8.0-1021-aws
Arch: aarch64
CPU Name:
C++ Compiler: /usr/bin/c++
C++ Compiler Version: 11.4.0
C Compiler: /usr/bin/cc
C Compiler Version: 11.4.0
CMake Prefix Path: /usr/local;/usr;/;/usr/local/lib/python3.10/dist-packages/cmake/data;/usr/local;/usr/X11R6;/usr/pkg;/opt

### Relevant logs

_No response_

Contributor guide

Open the contributing guide

Research direction

Start with tpch_parquet.scala and rerun TPC-H queries q7, q9, q10, and q12 using the two provided Spark command lines on the ARM-based AWS instance. Compare the timings and configurations, especially the forced shuffled-hash-join setting, then identify a reproducible cause and demonstrate an improvement with before-and-after results.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, linux, scala, sql
Domain
performance
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.