Gluten shuffle data size is twice that of vanilla Spark shuffle data size, with celeborn as remote shuffe service
Open
enhancement
- Dominant language
- Scala
- Stars
- 1.6k
- Forks
- 657
- Avg merge
- 2d 21h
- Merged PRs (30d)
- 85
Description
### Description
**vanilla spark**
**gluten**
shuffle from aggregate after data union
### Gluten version
None
Contributor guide
Research direction
No files, tests, versions, or entry points are named. Start by reproducing the vanilla Spark versus Gluten comparison with Celeborn enabled and tracing the aggregate-after-union shuffle configuration; done means explaining or correcting the doubled shuffle data size and verifying the comparison.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- scala
- Domain
- data-engineering, distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100