Azure / Azure/azure-sdk-tools

[TypeSpec Authoring][Benchmark] Evaluate token usage and investigate the optimization ways

Open
#15,731 1 comment 0 reactions 0 assignees View on GitHub
needs-triage
Dominant language
C#
Stars
135
Forks
260
Avg merge
3d 1h
Merged PRs (30d)
143

Description

Parent epic: #15726

Evaluate token usage of the `azure-typespec-author` skill during benchmark runs to understand cost drivers and find optimization opportunities.

Suggested scope:
- Measure input vs output token counts per stimulus / per scenario category (basic/advanced dataplane, basic/advanced ARM).
- Break down where tokens are spent: system prompt, skill instructions, MCP tool definitions, tool-call results, conversation history.
- Identify outliers (stimuli with disproportionately high token usage) and root-cause them.
- Compare token usage across models and across skill versions to track regressions/improvements.
- Feed the data into the quality monitor dashboard (#15728) so token cost trends are visible alongside pass rate.

Contributor guide

Open the contributing guide

Research direction

Start by reading parent epic #15726 and the benchmark runs for the azure-typespec-author skill, then review the quality monitor dashboard work in #15728. Measure input and output tokens by stimulus and scenario, identify token sources and outliers, compare models and skill versions, and feed the resulting trends into the dashboard.

Written by the indexing model from the issue text.

Assessment

Domain
ai, observability
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.