apache / apache/beam

Add Hadoop MapReduce runner

Open
#17,991 0 comments 0 reactions 0 assignees View on GitHub
ideas mapreduce new feature P3 runners
Dominant language
Java
Stars
8.7k
Forks
4.7k
Avg merge
2d 2h
Merged PRs (30d)
205

Description

I think a MapReduce runner could be a good addition to Beam. It would allow users to smoothly "migrate" from MapReduce to Spark or Flink.

Of course, the MapReduce runner will run in batch mode (not stream).

Imported from Jira [BEAM-165](https://issues.apache.org/jira/browse/BEAM-165). Original Jira may contain additional context.
Reported by: jbonofre.

Contributor guide

Open the contributing guide

Research direction

The issue names no files, tests, or entry points. Start by reviewing the linked BEAM-165 Jira ticket for its additional context, then determine the scope of a Hadoop MapReduce runner and its batch-only behavior. Done means Beam includes a usable runner that supports migration from MapReduce to Spark or Flink.

Written by the indexing model from the issue text.

Assessment

Tech stack
hadoop
Domain
data-engineering, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.