microsoft / microsoft/Power-CAT-Copilot-Studio-Kit
Add iterative execution of test cases for statistically relevant results
Open
@psimolin is already working on this.
Since Aug 13, 2025.
enhancement
- Dominant language
- TypeScript
- Stars
- 450
- Forks
- 95
- Avg merge
- 26m
- Merged PRs (30d)
- 5
Description
Because genAIi models are non-deterministic, a single execution of a test case will not produce a statistically significant result.
The single-run approach of test sets is likely to produce false positive/negatives and "flaky" test sets.
Adding a feature to iteratively run the test sets for a configured number times and report aggregated (average-based) results would significantly increase the accuracy of the evals.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.