apache / apache/gluten

Gluten shuffle data size is twice that of vanilla Spark shuffle data size, with celeborn as remote shuffe service

Open
#11,003 8 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Scala
Stars
1.6k
Forks
657
Avg merge
2d 21h
Merged PRs (30d)
85

Description

### Description

**vanilla spark**

Image

Image

**gluten**

Image

Image

shuffle from aggregate after data union

Image

### Gluten version

None

Contributor guide

Open the contributing guide

Research direction

No files, tests, versions, or entry points are named. Start by reproducing the vanilla Spark versus Gluten comparison with Celeborn enabled and tracing the aggregate-after-union shuffle configuration; done means explaining or correcting the doubled shuffle data size and verifying the comparison.

Written by the indexing model from the issue text.

Assessment

Tech stack
scala
Domain
data-engineering, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.