apache / apache/uniffle

[FEATURE] Support stage recompute for Tez clients

Open
#1,038 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
454
Forks
172
Avg merge
5d 17h
Merged PRs (30d)
5

Description

### Code of Conduct

- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

### Search before asking

- [X] I have searched in the [issues](https://github.com/apache/incubator-uniffle/issues?q=is%3Aissue) and found no similar issues.

### Describe the feature

For tez without remote shuffle server, when task recieve too many fetch exception from downstream task, this task will be recomputed.
With remote shuffle server, if some shuffle server is down, will loss data. Because data have been merged from many upstream task, so we must recompute upstream vertex.

This issue is the problem 2.d which is described in #1011

### Motivation

_No response_

### Describe the solution

_No response_

### Additional context

_No response_

### Are you willing to submit PR?

- [X] Yes I am willing to submit a PR!

Contributor guide

Open the contributing guide

Research direction

Start by reading issue #1011, especially problem 2.d, and trace how Tez clients handle downstream fetch exceptions with and without the remote shuffle server. Define the required stage-recompute behavior when a shuffle server is unavailable, then identify the affected entry points and tests. Done means upstream vertices can be recomputed without losing merged shuffle data.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.