jongalloway / jongalloway/dotnet-mcp
A/B testing: validate token savings heuristics against real LLM token counts
Nobody has claimed this yet.
- Dominant language
- C#
- Stars
- 36
- Forks
- 3
- PR merge metrics
- No merged PRs in 30d
Description
Summary
The token savings estimation system (TokenSavingsEstimator, TokenizerApproximation, ModelFamily) currently uses heuristic approximations — chars-per-token ratios and baseline scale factors — to estimate token counts. These values are educated guesses based on published tokenizer research, but they have not been validated against real-world MCP interactions.
We should run structured A/B comparisons to capture actual token counts and validate (or tune) these heuristics.
What needs validation
| Heuristic | Current Value (Unknown/default) | Notes |
|---|---|---|
| Prose chars-per-token | 4.0 | Based on GPT-4 BPE averages |
| JSON chars-per-token | 3.2 | Structured text tokenizes denser |
| Code chars-per-token | 3.5 | Symbols + identifiers |
| Baseline scale factor | 1.0 (varies by model family) | Multiplier on baseline overhead |
| Content-kind detection | density threshold 4% | {};() punctuation ratio for Code vs Prose |
Model-family-specific ratios also need validation (e.g., Claude Haiku at 3.8 prose, GPT-4o at 3.9, etc.).
Proposed approach
Phase 1: Instrumentation
- Add opt-in telemetry that captures actual token counts from LLM API responses alongside our heuristic estimates
- Log paired data:
(heuristic_estimate, actual_tokens, model_id, content_kind, content_length) - Store in a local JSONL file (privacy-first — no external telemetry without consent)
Phase 2: Data collection
- Run a representative set of MCP workflows (project creation, package management, build+test, template discovery) across multiple model families
- Capture at least ~100 paired observations per model family for statistical significance
- Include diverse content types: short prompts, long tool responses, JSON payloads, code output
Phase 3: Analysis & calibration
- Compute per-model-family mean absolute error (MAE) and mean absolute percentage error (MAPE)
- Fit updated chars-per-token ratios via least-squares regression on actual data
- Validate the
ContentKinddetection accuracy (confusion matrix: Prose/Json/Code) - Determine if
BaselineScaleFactorvalues track reality or need restructuring - Publish calibrated profile as
v2(keepingv1as fallback)
Phase 4: Ongoing validation
- Add a CI smoke test that compares heuristic estimates to a frozen set of known-good pairs
- Consider a
--calibratemode that auto-tunes from collected data
Success criteria
- MAPE < 15% across all model families for MCP token estimates
- MAPE < 25% for baseline estimates (inherently noisier due to prompt engineering variance)
ContentKinddetection accuracy > 90%- At least 3 model families validated (OpenAI GPT-4o, Claude Sonnet, Gemini Pro)
Related code
DotNetMcp/Telemetry/TokenSavingsEstimator.cs— core estimation logicDotNetMcp/Telemetry/TokenSavingsModels.cs—ModelFamily,ContentKind,TokenizerApproximationDotNetMcp.Tests/Telemetry/TokenSavingsEstimatorTests.cs— current test coverage
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with DotNetMcp/Telemetry/TokenSavingsEstimator.cs and TokenSavingsModels.cs, then review DotNetMcp.Tests/Telemetry/TokenSavingsEstimatorTests.cs to understand the existing heuristics and coverage. Define how opt-in paired observations are recorded in local JSONL and how representative MCP workflows provide actual counts. Done means calibrated results cover at least three model families, meet the stated MAPE and ContentKind accuracy targets, and include regression coverage for frozen known-good pairs.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- csharp
- Domain
- data, observability-sre, testing
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100