[TypeSpec Authoring] Add Vally assessment benchmarks
- Dominant language
- C#
- Stars
- 135
- Forks
- 260
- Avg merge
- 3d 1h
- Merged PRs (30d)
- 144
Description
## Goal
Add a Vally benchmark suite that measures TypeSpec assessment routing and report quality across all five dimensions.
## Scope
- Add trigger and neighboring-skill anti-trigger cases for the assessment skill.
- Add curated capability cases for semantic intent, REST breaking changes, downstream SDK breaking changes, Azure compliance, and documentation.
- Include positive, no-impact, ambiguous, and blocked-evidence cases.
- Grade source-link accuracy, finding classification, intent-to-impact links, incomplete-analysis status, and absence of unsupported findings.
- Use hermetic fixtures for the PR gate; keep any live benchmark in a separate serialized tier.
- Record per-case scores, duration, and model so benchmark changes are comparable over time.
## Acceptance criteria
- The benchmark follows the repository Vally eval-authoring conventions and runs from the existing skill-eval pipeline.
- Routing coverage includes at least three triggers and three meaningful anti-triggers with competing skills mounted where applicable.
- Every report dimension has representative regression cases and explicit graders.
- Known no-impact cases penalize false positives, and blocked cases cannot pass as safe.
- Results identify the failed scenario and grader and can be compared with a checked-in baseline.
- Focused Vally evaluation passes locally and in CI.
Contributor guide
Research direction
Start by reading the repository's Vally eval-authoring conventions and the existing skill-eval pipeline. Define the hermetic fixtures, routing cases, capability cases, graders, and checked-in baseline around those conventions; done means focused evaluation passes locally and in CI, results identify failed scenarios and graders, and live benchmarks remain a separate serialized tier.
Written by the indexing model from the issue text.
Assessment
- Domain
- testing, tooling
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100