microsoft / microsoft/hve-core
feat(agents): add end-to-end review-loop behavior eval for orientation-first Scenario D
- Dominant language
- Python
- Stars
- 1.5k
- Forks
- 301
- Avg merge
- 3d 3h
- Merged PRs (30d)
- 92
Description
Tracks WI-02 from the orientation-first review integration epic (#2090).
An end-to-end behavioral evaluation is needed to validate that the full Scenario D loop — orientation → dispatch board → walk-back → emission — functions correctly against a representative sample diff.
**Acceptance criteria:**
- A sample diff fixture is defined under `.copilot-tracking/` or a dedicated eval directory.
- The eval exercises each loop stage: orientation floor emits a narrative + dispatch board, walk-back resolves a finding, emission produces structured output.
- Eval results are documented or captured in a way that confirms expected behavior.
- Any behavioral discrepancies uncovered are logged as follow-on issues.
**References:**
- Parent epic: #2090
- Affected agents: `.github/agents/coding-standards/code-review.agent.md`, `code-review-walkback.agent.md`, `code-review-explainer.agent.md`
Related to #2090
> Generated by [Issue Triage](https://github.com/microsoft/hve-core/actions/runs/27888664228) for issue #2090 · 70.8 AIC · ⌖ 13.1 AIC · ⊞ 28.5K · [◷](https://github.com/search?q=repo%3Amicrosoft%2Fhve-core+is%3Aissue+%22gh-aw-workflow-call-id%3A+microsoft%2Fhve-core%2Fissue-triage%22&type=issues)
Contributor guide
Research direction
Read .github/agents/coding-standards/code-review.agent.md, code-review-walkback.agent.md, and code-review-explainer.agent.md, then review parent epic #2090 for the intended Scenario D flow. Define a representative sample diff and an evaluation path under .copilot-tracking/ or a dedicated eval directory. Done means orientation, dispatch, walk-back, and emission are exercised with expected results captured, and discrepancies are logged as follow-on issues.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, testing
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 52/100