ProjectTech4DevAI / ProjectTech4DevAI/kaapi-backend
Monitoring: Automate alerts and dashboards
@vprashrex ya está trabajando en esto.
Desde el 4/9/2026.
- Lenguaje dominante
- Python
- Estrellas
- 18
- Forks
- 10
- Merge medio
- 2 d 20 h
- PR fusionados (30 d)
- 14
Descripción
Is your feature request related to a problem?
We lack an ongoing monitoring and observability structure for Kaapi, leading to manual checks that can be inefficient and error-prone.
Describe the solution you'd like
- Automate breach notifications (memory, CloudWatch alarms) to Discord.
- Implement eval-run load monitoring in two stages:
- Stage 1: Sentry HTTP endpoint graphs on eval-run endpoints for run density.
- Stage 2: Live Celery queue backlog alerts.
- Establish app-level monitoring: endpoint hit counts, slow requests, waterfall dig-ins; create a custom Sentry dashboard.
- Create two dashboards: one for AWS (infra) and one for Sentry (app-level), to be reviewed in the weekly call.
- Weekly reporting should focus on red flags/odd numbers, skipping consistently green metrics.
Actions
- Propose feature-wise/app-level metrics (latency, etc.) for the Sentry dashboard.
- Set up Discord breach notifications.
- Verify staging Container Insights and configure staging alerts.
Related: ProjectTech4DevAI/kaapi-backend#1052, ProjectTech4DevAI/kaapi-backend#1005, ProjectTech4DevAI/kaapi-backend#825
Original issue
Context
Establishing the ongoing monitoring/observability structure for Kaapi. Akhilesh's health-check doc is taken as the infra baseline checklist. Direction is to automate rather than manually check.
Scope
- Push breach notifications (memory, CloudWatch alarms) to Discord instead of manual checks.
- Eval-run load monitoring:
stage 1 = Sentry HTTP endpoint graphs on eval-run endpoints as a proxy for run density;
stage 2 = live Celery queue backlog alerts. No LLM-call rate limits yet — deliberately deferred until scale demands it. - App-level monitoring: endpoint hit counts, slow requests, waterfall dig-ins; build a custom Sentry dashboard.
- Target: two dashboards — one AWS (infra), one Sentry (app-level) — reviewed in the weekly call.
- Weekly reporting convention: share only red flags / odd numbers; skip steadily-green metrics.
Actions
- Team to propose feature-wise / app-level metrics (latency etc.) for the Sentry dashboard.
- Set up Discord breach notifications.
- Verify staging Container Insights and configure staging alerts.
Related: ProjectTech4DevAI/kaapi-backend#1052, ProjectTech4DevAI/kaapi-backend#1005, ProjectTech4DevAI/kaapi-backend#825
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Evaluación
Este issue todavía no se ha evaluado.