dmwm / dmwm/CMSSpark

Evaluate to run yum update cern-hadoop-config in each Spark script

Open
#150 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
4
Forks
1
PR merge metrics
No merged PRs in 30d

Description

CERN Hadoop team started to change their config very frequently compared to the last years. Our Spark and Sqoop jobs fail because of non-updated config which includes connection parameters and Hadoop/Spark node urls.

We need to evaluate running yum update `cern-hadoop-config` before each Spark job running on Kubernetes pod.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the Spark job scripts and their Kubernetes pod entry points, then inspect how configuration packages are currently installed or refreshed. The work is done when there is a validated decision about running yum update cern-hadoop-config before each job and, if approved, the relevant jobs reliably apply it without breaking execution.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, spark
Domain
data-engineering, infrastructure
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.