github / github/github-mcp-server

get_diff / get_files: JSON serialization inflates diff payload far beyond token limits

Open
#2,242 1 comment 0 reactions 0 assignees View on GitHub
request ai review
Dominant language
Go
Stars
33k
Forks
5k
Avg merge
2d 1h
Merged PRs (30d)
52

Description

## Problem

`pull_request_read` methods `get_diff` and `get_files` return diff content as JSON-encoded strings. A unified diff is already a text format, but wrapping it in a JSON string escapes every `\n`, `\"`, `\t`, etc. — inflating the payload **several times** over the raw text size.

### Real-world example

A PR with **2,920 lines** of diff (~73 KB of raw text) produces a **125 KB+** JSON response, exceeding the MCP token limit. The tool returns an error instead of the diff.

This is a ~1,700-line PR (additions + deletions) — not unusually large. It includes some generated files (`src/generated/graphql.ts`), but even without them the JSON overhead makes moderate PRs hit the limit.

### Why this matters

- The token limit is hit not because the diff is large, but because JSON serialization inflates it
- `get_files` has the same problem — each file's `patch` field is a JSON-encoded diff string
- This forces users to fall back to `gh pr diff` via shell, losing the benefit of MCP

## Suggestion

Return diff content as raw text rather than a JSON-wrapped string, or use a more efficient serialization that doesn't escape every newline. This would let `get_diff` handle PRs several times larger than it can today with no other changes.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.