ProjectTech4DevAI / ProjectTech4DevAI/kaapi-backend

Monitoring: Automate alerts and dashboards

Open
#1,144 0 comments 0 reactions 1 assignee View on GitHub

@vprashrex is already working on this.

Since Sep 4, 2026.

Dominant language
Python
Stars
18
Forks
10
Avg merge
2d 20h
Merged PRs (30d)
14

Description

Is your feature request related to a problem?
We lack an ongoing monitoring and observability structure for Kaapi, leading to manual checks that can be inefficient and error-prone.

Describe the solution you'd like

  • Automate breach notifications (memory, CloudWatch alarms) to Discord.
  • Implement eval-run load monitoring in two stages:
    • Stage 1: Sentry HTTP endpoint graphs on eval-run endpoints for run density.
    • Stage 2: Live Celery queue backlog alerts.
  • Establish app-level monitoring: endpoint hit counts, slow requests, waterfall dig-ins; create a custom Sentry dashboard.
  • Create two dashboards: one for AWS (infra) and one for Sentry (app-level), to be reviewed in the weekly call.
  • Weekly reporting should focus on red flags/odd numbers, skipping consistently green metrics.

Actions

  • Propose feature-wise/app-level metrics (latency, etc.) for the Sentry dashboard.
  • Set up Discord breach notifications.
  • Verify staging Container Insights and configure staging alerts.

Related: ProjectTech4DevAI/kaapi-backend#1052, ProjectTech4DevAI/kaapi-backend#1005, ProjectTech4DevAI/kaapi-backend#825

Original issue

Context

Establishing the ongoing monitoring/observability structure for Kaapi. Akhilesh's health-check doc is taken as the infra baseline checklist. Direction is to automate rather than manually check.

Scope

  • Push breach notifications (memory, CloudWatch alarms) to Discord instead of manual checks.
  • Eval-run load monitoring:
    stage 1 = Sentry HTTP endpoint graphs on eval-run endpoints as a proxy for run density;
    stage 2 = live Celery queue backlog alerts. No LLM-call rate limits yet — deliberately deferred until scale demands it.
  • App-level monitoring: endpoint hit counts, slow requests, waterfall dig-ins; build a custom Sentry dashboard.
  • Target: two dashboards — one AWS (infra), one Sentry (app-level) — reviewed in the weekly call.
  • Weekly reporting convention: share only red flags / odd numbers; skip steadily-green metrics.

Actions

  • Team to propose feature-wise / app-level metrics (latency etc.) for the Sentry dashboard.
  • Set up Discord breach notifications.
  • Verify staging Container Insights and configure staging alerts.

Related: ProjectTech4DevAI/kaapi-backend#1052, ProjectTech4DevAI/kaapi-backend#1005, ProjectTech4DevAI/kaapi-backend#825

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.