Agent-Hellboy / Agent-Hellboy/mcp-runtime

Add Prometheus latency metrics for Sentinel services and MCP server calls

Đang mở
#285 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Go
Star
6
Fork
1
Merge trung bình
11 giờ 33 phút
Pull request đã merge (30 ngày)
13

Mô tả

## Problem

Prometheus currently scrapes the Sentinel services and ClickHouse, but it does not expose request-duration histograms for MCP operations. The live cluster only has scrape/runtime metrics in Prometheus, while request latency is currently available indirectly from ClickHouse audit payloads via `latency_ms`.

That is useful for historical analysis, but it is not enough for live Grafana dashboards, SLOs, alerting, or fast operational debugging.

## What to add

Add Prometheus metrics for MCP request latency and throughput, covering both Sentinel services and MCP server/gateway paths, at minimum:
- `tools/call`
- prompt calls
- resource reads
- any other MCP request path the gateway/service actively serves

Suggested metric shape:
- `mcp_request_duration_seconds`
- `mcp_request_total`
- `mcp_request_errors_total`

Use low-cardinality labels only, such as:
- `service`
- `operation`
- `status`
- `server`

Avoid labels that can explode cardinality, like user IDs, session IDs, prompt names, or raw paths.

## Acceptance criteria

- Prometheus exposes histogram buckets for MCP request latency.
- Grafana can compute p50/p95/p99 from live Prometheus data.
- Metrics cover tool calls, prompt calls, resource reads, and the Sentinel service path where applicable.
- Tests or smoke checks verify the metric names exist.

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.