ProjectTech4DevAI / ProjectTech4DevAI/kaapi-backend

Monitoring: Automate alerts and dashboards

Abierto
#1,144 0 comentarios 0 reacciones 1 asignado Ver en GitHub

@vprashrex ya está trabajando en esto.

Desde el 4/9/2026.

Lenguaje dominante
Python
Estrellas
18
Forks
10
Merge medio
2 d 20 h
PR fusionados (30 d)
14

Descripción

Is your feature request related to a problem?
We lack an ongoing monitoring and observability structure for Kaapi, leading to manual checks that can be inefficient and error-prone.

Describe the solution you'd like

  • Automate breach notifications (memory, CloudWatch alarms) to Discord.
  • Implement eval-run load monitoring in two stages:
    • Stage 1: Sentry HTTP endpoint graphs on eval-run endpoints for run density.
    • Stage 2: Live Celery queue backlog alerts.
  • Establish app-level monitoring: endpoint hit counts, slow requests, waterfall dig-ins; create a custom Sentry dashboard.
  • Create two dashboards: one for AWS (infra) and one for Sentry (app-level), to be reviewed in the weekly call.
  • Weekly reporting should focus on red flags/odd numbers, skipping consistently green metrics.

Actions

  • Propose feature-wise/app-level metrics (latency, etc.) for the Sentry dashboard.
  • Set up Discord breach notifications.
  • Verify staging Container Insights and configure staging alerts.

Related: ProjectTech4DevAI/kaapi-backend#1052, ProjectTech4DevAI/kaapi-backend#1005, ProjectTech4DevAI/kaapi-backend#825

Original issue

Context

Establishing the ongoing monitoring/observability structure for Kaapi. Akhilesh's health-check doc is taken as the infra baseline checklist. Direction is to automate rather than manually check.

Scope

  • Push breach notifications (memory, CloudWatch alarms) to Discord instead of manual checks.
  • Eval-run load monitoring:
    stage 1 = Sentry HTTP endpoint graphs on eval-run endpoints as a proxy for run density;
    stage 2 = live Celery queue backlog alerts. No LLM-call rate limits yet — deliberately deferred until scale demands it.
  • App-level monitoring: endpoint hit counts, slow requests, waterfall dig-ins; build a custom Sentry dashboard.
  • Target: two dashboards — one AWS (infra), one Sentry (app-level) — reviewed in the weekly call.
  • Weekly reporting convention: share only red flags / odd numbers; skip steadily-green metrics.

Actions

  • Team to propose feature-wise / app-level metrics (latency etc.) for the Sentry dashboard.
  • Set up Discord breach notifications.
  • Verify staging Container Insights and configure staging alerts.

Related: ProjectTech4DevAI/kaapi-backend#1052, ProjectTech4DevAI/kaapi-backend#1005, ProjectTech4DevAI/kaapi-backend#825

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.