Agent-Hellboy / Agent-Hellboy/mcp-runtime
Add Prometheus latency metrics for Sentinel services and MCP server calls
- Lingua principale
- Go
- Stelle
- 6
- Fork
- 1
- Merge medio
- 11h 33m
- PR unite (30g)
- 13
Descrizione
## Problem
Prometheus currently scrapes the Sentinel services and ClickHouse, but it does not expose request-duration histograms for MCP operations. The live cluster only has scrape/runtime metrics in Prometheus, while request latency is currently available indirectly from ClickHouse audit payloads via `latency_ms`.
That is useful for historical analysis, but it is not enough for live Grafana dashboards, SLOs, alerting, or fast operational debugging.
## What to add
Add Prometheus metrics for MCP request latency and throughput, covering both Sentinel services and MCP server/gateway paths, at minimum:
- `tools/call`
- prompt calls
- resource reads
- any other MCP request path the gateway/service actively serves
Suggested metric shape:
- `mcp_request_duration_seconds`
- `mcp_request_total`
- `mcp_request_errors_total`
Use low-cardinality labels only, such as:
- `service`
- `operation`
- `status`
- `server`
Avoid labels that can explode cardinality, like user IDs, session IDs, prompt names, or raw paths.
## Acceptance criteria
- Prometheus exposes histogram buckets for MCP request latency.
- Grafana can compute p50/p95/p99 from live Prometheus data.
- Metrics cover tool calls, prompt calls, resource reads, and the Sentinel service path where applicable.
- Tests or smoke checks verify the metric names exist.
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.