Azure / Azure/azure-sdk-tools

[TypeSpec Authoring] Add Vally assessment benchmarks

Open
#16,830 0 comments 0 reactions 0 assignees View on GitHub
AzSDK Tools Agent dev inner loop needs-triage
Dominant language
C#
Stars
135
Forks
260
Avg merge
3d 1h
Merged PRs (30d)
144

Description

## Goal

Add a Vally benchmark suite that measures TypeSpec assessment routing and report quality across all five dimensions.

## Scope

- Add trigger and neighboring-skill anti-trigger cases for the assessment skill.
- Add curated capability cases for semantic intent, REST breaking changes, downstream SDK breaking changes, Azure compliance, and documentation.
- Include positive, no-impact, ambiguous, and blocked-evidence cases.
- Grade source-link accuracy, finding classification, intent-to-impact links, incomplete-analysis status, and absence of unsupported findings.
- Use hermetic fixtures for the PR gate; keep any live benchmark in a separate serialized tier.
- Record per-case scores, duration, and model so benchmark changes are comparable over time.

## Acceptance criteria

- The benchmark follows the repository Vally eval-authoring conventions and runs from the existing skill-eval pipeline.
- Routing coverage includes at least three triggers and three meaningful anti-triggers with competing skills mounted where applicable.
- Every report dimension has representative regression cases and explicit graders.
- Known no-impact cases penalize false positives, and blocked cases cannot pass as safe.
- Results identify the failed scenario and grader and can be compared with a checked-in baseline.
- Focused Vally evaluation passes locally and in CI.

Contributor guide

Open the contributing guide

Research direction

Start by reading the repository's Vally eval-authoring conventions and the existing skill-eval pipeline. Define the hermetic fixtures, routing cases, capability cases, graders, and checked-in baseline around those conventions; done means focused evaluation passes locally and in CI, results identify failed scenarios and graders, and live benchmarks remain a separate serialized tier.

Written by the indexing model from the issue text.

Assessment

Domain
testing, tooling
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.