microsoft / microsoft/hve-core

feat(agents): add end-to-end review-loop behavior eval for orientation-first Scenario D

Open
#2,093 0 comments 0 reactions 0 assignees View on GitHub
agents feature needs-triage
Dominant language
Python
Stars
1.5k
Forks
301
Avg merge
3d 3h
Merged PRs (30d)
92

Description

Tracks WI-02 from the orientation-first review integration epic (#2090).

An end-to-end behavioral evaluation is needed to validate that the full Scenario D loop — orientation → dispatch board → walk-back → emission — functions correctly against a representative sample diff.

**Acceptance criteria:**
- A sample diff fixture is defined under `.copilot-tracking/` or a dedicated eval directory.
- The eval exercises each loop stage: orientation floor emits a narrative + dispatch board, walk-back resolves a finding, emission produces structured output.
- Eval results are documented or captured in a way that confirms expected behavior.
- Any behavioral discrepancies uncovered are logged as follow-on issues.

**References:**
- Parent epic: #2090
- Affected agents: `.github/agents/coding-standards/code-review.agent.md`, `code-review-walkback.agent.md`, `code-review-explainer.agent.md`
Related to #2090

> Generated by [Issue Triage](https://github.com/microsoft/hve-core/actions/runs/27888664228) for issue #2090 · 70.8 AIC · ⌖ 13.1 AIC · ⊞ 28.5K · [◷](https://github.com/search?q=repo%3Amicrosoft%2Fhve-core+is%3Aissue+%22gh-aw-workflow-call-id%3A+microsoft%2Fhve-core%2Fissue-triage%22&type=issues)

Contributor guide

Open the contributing guide

Research direction

Read .github/agents/coding-standards/code-review.agent.md, code-review-walkback.agent.md, and code-review-explainer.agent.md, then review parent epic #2090 for the intended Scenario D flow. Define a representative sample diff and an evaluation path under .copilot-tracking/ or a dedicated eval directory. Done means orientation, dispatch, walk-back, and emission are exercised with expected results captured, and discrepancies are logged as follow-on issues.

Written by the indexing model from the issue text.

Assessment

Domain
ai, testing
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.