ISISComputingGroup / ISISComputingGroup/IBEX

System & server health check (independent of GUI & nagios)

オープン
#3,752 コメント 17 件 リアクション 0 件 担当者 0 名 GitHub で見る
Friday
主要言語
言語のデータがありません
スター
6
フォーク
2
平均マージ
16時間 40分
マージ済み PR(30日)
2

説明

As a VESUVIO instrument scientist I want to know if either my block or instrument archiver is not running so that I can take remedial action.

This information in cycle is alerted on by Nagios but the instrument scientist would like to know.

Preferably something similar to the error users see when the block server is not up.

NB Check should be that it is healthy not the process is around.

## Acceptance Criteria
1. Finish defining the acceptance criteria with what we expect
2. Consider the distribution methods
3. Nagios may be enough on it's own

## Notes

Per discussion in comments, this would ideally be a check independent of the GUI, which would then be usable by other components (e.g. potentially the web dashboard, genie_python). We may wish to internally re-use some of the nagios infrastructure, but this should be ultimately scientist-facing which nagios is not.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

まず、既存の Nagios checks と、block server が利用できない場合に表示される GUI メッセージを確認し、次に web dashboard または genie_python で check を再利用できるか検討してください。“healthy” が何を意味するか、また Nagios だけで十分かどうかを含め、受け入れ基準と配布方法を定義してください。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
python
領域
backend-api-design, observability-sre
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。