Agent-Hellboy / Agent-Hellboy/mcp-runtime

Add Prometheus latency metrics for Sentinel services and MCP server calls

未关闭
#285 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Go
星标
6
派生
1
平均合并
11 小时 33 分钟
30 天内合并 PR
13

描述

## Problem

Prometheus currently scrapes the Sentinel services and ClickHouse, but it does not expose request-duration histograms for MCP operations. The live cluster only has scrape/runtime metrics in Prometheus, while request latency is currently available indirectly from ClickHouse audit payloads via `latency_ms`.

That is useful for historical analysis, but it is not enough for live Grafana dashboards, SLOs, alerting, or fast operational debugging.

## What to add

Add Prometheus metrics for MCP request latency and throughput, covering both Sentinel services and MCP server/gateway paths, at minimum:
- `tools/call`
- prompt calls
- resource reads
- any other MCP request path the gateway/service actively serves

Suggested metric shape:
- `mcp_request_duration_seconds`
- `mcp_request_total`
- `mcp_request_errors_total`

Use low-cardinality labels only, such as:
- `service`
- `operation`
- `status`
- `server`

Avoid labels that can explode cardinality, like user IDs, session IDs, prompt names, or raw paths.

## Acceptance criteria

- Prometheus exposes histogram buckets for MCP request latency.
- Grafana can compute p50/p95/p99 from live Prometheus data.
- Metrics cover tool calls, prompt calls, resource reads, and the Sentinel service path where applicable.
- Tests or smoke checks verify the metric names exist.

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。