Agent-Hellboy / Agent-Hellboy/mcp-runtime

feat(mcp-gateway): per-tool Prometheus golden metrics

Abierto
#300 0 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Go
Estrellas
6
Forks
1
Merge medio
11 h 33 min
PR fusionados (30 d)
13

Descripción

## Summary

The gateway currently only exposes policy-reload metrics. Add per-tool call observability matching what Linkerd exposes per route — success rate, RPS, and latency histograms broken down by tool name.

## Proposed metrics

```
mcp_gateway_tool_call_duration_seconds{tool, decision, server} // Histogram
mcp_gateway_tool_call_total{tool, decision, reason, server} // Counter
```

- `decision` — `allow` or `deny`
- `reason` — denial reason (`no_matching_grant`, `tool_denied`, etc.) or `allowed`
- `server` — from `policy.Server.Name`
- Histogram buckets covering slow tools: `5ms … 60s`

## Why

The `Exchange` struct already carries tool name, latency, decision, and server at audit time — wiring metrics is a small addition to `emitAuditFromExchange`. The existing `/metrics` endpoint and Grafana dashboard pick them up automatically.

This enables per-tool success rate and p50/p95/p99 latency heatmaps with zero dashboard changes.

## Scope

- `services/mcp-gateway/metrics.go` — new histogram + counter + `recordToolCall` helper
- `services/mcp-gateway/proxy.go` — call `recordToolCall` from `emitAuditFromExchange`
- `services/mcp-gateway/main.go` — already has `/metrics`; ensure `promhttp` is registered

Only record when `toolName != ""` — list/ping calls should not pollute the tool histogram with empty labels.

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.