Investigate how we can better account for variance between runs
Open
Evals
- Dominant language
- C#
- Stars
- 135
- Forks
- 260
- Avg merge
- 3d 1h
- Merged PRs (30d)
- 143
Description
Given non-deterministic responses by the model/agent, it can be hard to determine whether a decrease in a metric is due to changes made to the agents/prompts or just due to variance/noise.
Contributor guide
Research direction
No files, tests, or entry points are named. Start by locating the code that records or compares agent and prompt metrics, then define how variance should be measured and what evidence would show that run-to-run noise is accounted for.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- csharp
- Domain
- ai, analytics
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100