apache / apache/gluten

Almost 10 times Hash Agg spill memory than vanilla Spark

Open
#9,395 0 comments 0 reactions 0 assignees View on GitHub
bug triage
Dominant language
Scala
Stars
1.6k
Forks
657
Avg merge
2d 21h
Merged PRs (30d)
85

Description

### Backend

VL (Velox)

### Bug description

we found that for some case, Hash Agg hash huge memory spilled, and the stage with huge memory spill is running slower than vanilla Spark

### Gluten version

_No response_

### Spark version

None

### Spark configurations

_No response_

### System information

_No response_

### Relevant logs

```bash

```

Contributor guide

Open the contributing guide

Research direction

Start by obtaining a reproducible Velox hash-aggregation case and the missing Gluten, Spark, configuration, system, and log details. Compare its spill memory and stage runtime with vanilla Spark; done means the excess spill and slowdown are reproduced and their cause is identified.

Written by the indexing model from the issue text.

Assessment

Domain
backend, data
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.