apache / apache/uniffle

[FEATURE] Group by the small segments on getting local shuffle data

Open
#2,510 0 comments 0 reactions 0 assignees View on GitHub
good first issue
Dominant language
Java
Stars
454
Forks
172
Avg merge
5d 17h
Merged PRs (30d)
5

Description

### Code of Conduct

- [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

### Search before asking

- [x] I have searched in the [issues](https://github.com/apache/incubator-uniffle/issues?q=is%3Aissue) and found no similar issues.

### Describe the feature

If hitting the AQE optimization, the partial segments will be small when filtering by mapId. For this case, we could group the small segments into one rpc to reduce the overhead

### Motivation

_No response_

### Describe the solution

_No response_

### Additional context

_No response_

### Are you willing to submit PR?

- [ ] Yes I am willing to submit a PR!

Contributor guide

Open the contributing guide

Research direction

Start by tracing the local-shuffle-data retrieval path and its AQE/mapId filtering behavior; the issue names no files or tests. Confirm the existing RPC boundaries, then treat completion as grouping small partial segments into one RPC without changing the retrieved data.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, spark
Domain
distributed-systems, performance
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.