ISISComputingGroup / ISISComputingGroup/IBEX

System & server health check (independent of GUI & nagios)

Offen
#3,752 17 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Friday
Vorherrschende Sprache
Keine Sprachdaten
Sterne
6
Forks
2
Ø Merge
16 Std. 40 Min.
Gemergte PRs (30 T.)
2

Beschreibung

As a VESUVIO instrument scientist I want to know if either my block or instrument archiver is not running so that I can take remedial action.

This information in cycle is alerted on by Nagios but the instrument scientist would like to know.

Preferably something similar to the error users see when the block server is not up.

NB Check should be that it is healthy not the process is around.

## Acceptance Criteria
1. Finish defining the acceptance criteria with what we expect
2. Consider the distribution methods
3. Nagios may be enough on it's own

## Notes

Per discussion in comments, this would ideally be a check independent of the GUI, which would then be usable by other components (e.g. potentially the web dashboard, genie_python). We may wish to internally re-use some of the nagios infrastructure, but this should be ultimately scientist-facing which nagios is not.

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Start by reviewing the existing Nagios checks and the GUI message shown when the block server is unavailable, then consider how the check could be reused by the web dashboard or genie_python. Define the acceptance criteria and distribution method, including what “healthy” means and whether Nagios alone is sufficient.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
python
Bereich
backend-api-design, observability-sre
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.