[TypeSpec Authoring][Benchmark] Evaluate token usage and investigate the optimization ways
- Dominant language
- C#
- Stars
- 135
- Forks
- 260
- Avg merge
- 3d 1h
- Merged PRs (30d)
- 143
Description
Parent epic: #15726
Evaluate token usage of the `azure-typespec-author` skill during benchmark runs to understand cost drivers and find optimization opportunities.
Suggested scope:
- Measure input vs output token counts per stimulus / per scenario category (basic/advanced dataplane, basic/advanced ARM).
- Break down where tokens are spent: system prompt, skill instructions, MCP tool definitions, tool-call results, conversation history.
- Identify outliers (stimuli with disproportionately high token usage) and root-cause them.
- Compare token usage across models and across skill versions to track regressions/improvements.
- Feed the data into the quality monitor dashboard (#15728) so token cost trends are visible alongside pass rate.
Contributor guide
Research direction
Start by reading parent epic #15726 and the benchmark runs for the azure-typespec-author skill, then review the quality monitor dashboard work in #15728. Measure input and output tokens by stimulus and scenario, identify token sources and outliers, compare models and skill versions, and feed the resulting trends into the dashboard.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, observability
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100