Agent-Hellboy / Agent-Hellboy/mcp-runtime

feat(mcp-gateway): per-tool Prometheus golden metrics

未關閉
#300 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Go
星號
6
分支
1
平均合併
11 小時 33 分鐘
30 天內合併 PR
13

描述

## Summary

The gateway currently only exposes policy-reload metrics. Add per-tool call observability matching what Linkerd exposes per route — success rate, RPS, and latency histograms broken down by tool name.

## Proposed metrics

```
mcp_gateway_tool_call_duration_seconds{tool, decision, server} // Histogram
mcp_gateway_tool_call_total{tool, decision, reason, server} // Counter
```

- `decision` — `allow` or `deny`
- `reason` — denial reason (`no_matching_grant`, `tool_denied`, etc.) or `allowed`
- `server` — from `policy.Server.Name`
- Histogram buckets covering slow tools: `5ms … 60s`

## Why

The `Exchange` struct already carries tool name, latency, decision, and server at audit time — wiring metrics is a small addition to `emitAuditFromExchange`. The existing `/metrics` endpoint and Grafana dashboard pick them up automatically.

This enables per-tool success rate and p50/p95/p99 latency heatmaps with zero dashboard changes.

## Scope

- `services/mcp-gateway/metrics.go` — new histogram + counter + `recordToolCall` helper
- `services/mcp-gateway/proxy.go` — call `recordToolCall` from `emitAuditFromExchange`
- `services/mcp-gateway/main.go` — already has `/metrics`; ensure `promhttp` is registered

Only record when `toolName != ""` — list/ping calls should not pollute the tool histogram with empty labels.

貢獻指南

這個儲存庫沒有索引到貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。