apache / apache/uniffle

[FEATURE] [MR] Fast fail job when too many fetch failed happen.

Open
#1,042 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
454
Forks
172
Avg merge
5d 17h
Merged PRs (30d)
5

Description

### Code of Conduct

- [X] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct)

### Search before asking

- [X] I have searched in the [issues](https://github.com/apache/incubator-uniffle/issues?q=is%3Aissue) and found no similar issues.

### Describe the feature

In #477 and #1011, job will resubmit upstream stage/vertex when many fetch failed. It is not useful for MR, because there are only 2 stage. For MR, we should fast fail the job.
But now MR can not fast fail. When the selected shuffle server is down or stuck, MR job fail until the number of task attempts is larger than max. It is very slow.
For now, shuffle does not send fetch failed event. We should fail the job according to the number of fetch failed event.

### Motivation

_No response_

### Describe the solution

_No response_

### Additional context

_No response_

### Are you willing to submit PR?

- [X] Yes I am willing to submit a PR!

Contributor guide

Open the contributing guide

Research direction

Read the resubmission behavior described in issues #477 and #1011, then trace how MapReduce handles fetch-failed events and task-attempt limits. Determine where fetch-failure counts can trigger job failure for MR, and verify that a selected shuffle server becoming unavailable causes fast failure instead of waiting for the maximum attempts.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data-engineering, distributed-systems
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.