ISISComputingGroup / ISISComputingGroup/IBEX
System & server health check (independent of GUI & nagios)
- Lenguaje dominante
- Sin datos de lenguaje
- Estrellas
- 6
- Forks
- 2
- Merge medio
- 16 h 40 min
- PR fusionados (30 d)
- 2
Descripción
As a VESUVIO instrument scientist I want to know if either my block or instrument archiver is not running so that I can take remedial action.
This information in cycle is alerted on by Nagios but the instrument scientist would like to know.
Preferably something similar to the error users see when the block server is not up.
NB Check should be that it is healthy not the process is around.
## Acceptance Criteria
1. Finish defining the acceptance criteria with what we expect
2. Consider the distribution methods
3. Nagios may be enough on it's own
## Notes
Per discussion in comments, this would ideally be a check independent of the GUI, which would then be usable by other components (e.g. potentially the web dashboard, genie_python). We may wish to internally re-use some of the nagios infrastructure, but this should be ultimately scientist-facing which nagios is not.
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
Empieza revisando los checks de Nagios existentes y el mensaje de la GUI que se muestra cuando el block server no está disponible; después, considera cómo podría reutilizarse el check desde el web dashboard o genie_python. Define los criterios de aceptación y el método de distribución, incluido qué significa “healthy” y si Nagios por sí solo es suficiente.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python
- Área
- backend-api-design, observability-sre
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Estancado
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 25/100