Agent-Hellboy / Agent-Hellboy/mcp-runtime

Add Prometheus latency metrics for Sentinel services and MCP server calls

Abierto
#285 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Go
Estrellas
6
Forks
1
Merge medio
11 h 33 min
PR fusionados (30 d)
13

Descripción

## Problem

Prometheus currently scrapes the Sentinel services and ClickHouse, but it does not expose request-duration histograms for MCP operations. The live cluster only has scrape/runtime metrics in Prometheus, while request latency is currently available indirectly from ClickHouse audit payloads via `latency_ms`.

That is useful for historical analysis, but it is not enough for live Grafana dashboards, SLOs, alerting, or fast operational debugging.

## What to add

Add Prometheus metrics for MCP request latency and throughput, covering both Sentinel services and MCP server/gateway paths, at minimum:
- `tools/call`
- prompt calls
- resource reads
- any other MCP request path the gateway/service actively serves

Suggested metric shape:
- `mcp_request_duration_seconds`
- `mcp_request_total`
- `mcp_request_errors_total`

Use low-cardinality labels only, such as:
- `service`
- `operation`
- `status`
- `server`

Avoid labels that can explode cardinality, like user IDs, session IDs, prompt names, or raw paths.

## Acceptance criteria

- Prometheus exposes histogram buckets for MCP request latency.
- Grafana can compute p50/p95/p99 from live Prometheus data.
- Metrics cover tool calls, prompt calls, resource reads, and the Sentinel service path where applicable.
- Tests or smoke checks verify the metric names exist.

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.