Azure / Azure/azure-sdk-tools

Investigate how we can better account for variance between runs

Open
#11,318 1 comment 0 reactions 0 assignees View on GitHub
Evals
Dominant language
C#
Stars
135
Forks
260
Avg merge
3d 1h
Merged PRs (30d)
143

Description

Given non-deterministic responses by the model/agent, it can be hard to determine whether a decrease in a metric is due to changes made to the agents/prompts or just due to variance/noise.

Contributor guide

Open the contributing guide

Research direction

No files, tests, or entry points are named. Start by locating the code that records or compares agent and prompt metrics, then define how variance should be measured and what evidence would show that run-to-run noise is accounted for.

Written by the indexing model from the issue text.

Assessment

Tech stack
csharp
Domain
ai, analytics
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.