Evaluate to run yum update cern-hadoop-config in each Spark script
- Dominant language
- Python
- Stars
- 4
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
CERN Hadoop team started to change their config very frequently compared to the last years. Our Spark and Sqoop jobs fail because of non-updated config which includes connection parameters and Hadoop/Spark node urls.
We need to evaluate running yum update `cern-hadoop-config` before each Spark job running on Kubernetes pod.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating the Spark job scripts and their Kubernetes pod entry points, then inspect how configuration packages are currently installed or refreshed. The work is done when there is a validated decision about running yum update cern-hadoop-config before each job and, if approved, the relevant jobs reliably apply it without breaking execution.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- kubernetes, spark
- Domain
- data-engineering, infrastructure
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100