microsoft / microsoft/Power-CAT-Copilot-Studio-Kit

Add iterative execution of test cases for statistically relevant results

Open
#277 1 comment 0 reactions 1 assignee View on GitHub

@psimolin is already working on this.

Since Aug 13, 2025.

enhancement
Dominant language
TypeScript
Stars
450
Forks
95
Avg merge
26m
Merged PRs (30d)
5

Description

Because genAIi models are non-deterministic, a single execution of a test case will not produce a statistically significant result.
The single-run approach of test sets is likely to produce false positive/negatives and "flaky" test sets.
Adding a feature to iteratively run the test sets for a configured number times and report aggregated (average-based) results would significantly increase the accuracy of the evals.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.