llnl / llnl/magpie

Support some sort of Monitoring Software

Open
#73 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

PotentialEnhancement
Dominant language
Shell
Stars
198
Forks
51
PR merge metrics
No merged PRs in 30d

Description

It would be great it we added something like ambari so that we could check the status of our nodes. In one case, all of my HBase region servers went down and my Spark job just hung around waiting for them to come back up (which they never did due to an unrelated issue). I had no idea that this occurred. It would be nice to be able to see what is happening all in one place.

I'm not sure what else is out that other than ambari at this point.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are identified in the issue. Start by surveying the repository's Hadoop and Spark workflow scripts and determine whether node-status monitoring fits the project; done would require an agreed monitoring approach that exposes failed nodes and stalled jobs in one place.

Written by the indexing model from the issue text.

Assessment

Tech stack
hadoop, shell, spark
Domain
distributed-systems, observability-sre
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.