github / github/github-mcp-server

Benchmark and improve tool-search ranking with indexed BM25

未關閉
#2,996 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
enhancement request ai review
主要語言
Go
星號
33k
分支
5k
平均合併
2 天 1 小時
30 天內合併 PR
52

描述

### Describe the feature or problem you’d like to solve

The GitHub MCP Server already exposes tool discovery/search functionality, but
there is no repeatable benchmark for measuring how reliably natural-language
queries retrieve the intended MCP tool.

As the tool inventory grows, a benchmark would make ranking changes measurable
and help prevent retrieval regressions.

This is separate from host-side deferred tool loading discussed in #1680. The
proposal only concerns ranking inside the server's existing tool-search
implementation.

### Proposed solution

Add a hand-labelled benchmark covering natural-language intents across the
server's major toolsets, then compare the current heuristic with an indexed
BM25 implementation.

A prototype benchmark contains 49 queries over 115 unique tools and produced:

| Strategy | Recall@1 | Recall@3 | MRR@10 | Query latency |
|---|---:|---:|---:|---:|
| Current heuristic | 71.4% | 81.6% | 0.792 | ~2.25 ms |
| Indexed BM25 | 71.4% | 87.8% | 0.802 | ~34 µs |
| Hybrid RRF | 73.5% | 87.8% | 0.823 | ~2.38 ms |

Indexed BM25 improved Recall@3 by 6.1 percentage points and was approximately
66x faster per query. The hybrid produced the strongest ranking quality.

Before submitting a PR, I would appreciate maintainer guidance on the preferred
scope:

1. Benchmark harness only
2. Benchmark plus indexed BM25
3. Benchmark plus a hybrid ranking experiment

### Example prompts or workflows

- "Find open issues assigned to me across repositories"
- "Read the files, reviews, and diff for a pull request"
- "Download logs for a failed workflow job"
- "Find exposed secrets detected in a repository"
- "Add an issue to a GitHub project"

### Additional context

The benchmark uses the complete current tool inventory and validates that every
labelled relevant tool exists. The prototype includes unit tests for indexing
tool names, descriptions, parameter names, and parameter descriptions.

貢獻指南

開啟貢獻指南

研究方向

Start by locating the server’s existing tool-search implementation and reviewing the prototype benchmark and its unit tests for indexing tool names, descriptions, parameter names, and parameter descriptions. Confirm the preferred scope with maintainers, then measure the selected ranking strategy against the 49-query, 115-tool benchmark and report retrieval quality and latency without regressions.

由索引模型根據 Issue 內容生成。

評估

技術堆疊
go
領域
backend-api-design, search
Issue 類型
功能
難度
5/5
預估耗時
一週以上
活躍度
活躍
描述清晰度
基本清楚
新手友好度
45/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。