microsoft / microsoft/GitHub-Copilot-for-Azure

Replace microsoft-foundry skill eval-results download-and-summarize with a script

Open
#2,536 1 comment 0 reactions 1 assignee Claimed by @tmeschter View on GitHub
microsoft-foundry skills untriaged
Dominant language
Python
Stars
250
Forks
204
Avg merge
1d 12h
Merged PRs (30d)
67

Description

## Summary

Copilot has identified a portion of a skill that is a good candidate for replacement with a script.

The candidate is the **eval-results download-and-summarize** flow in the `microsoft-foundry` skill (`foundry-agent/observe/references/analyze-results.md`) — a fixed SDK download → paginated list → dual-entry parse → persist → summary sequence that the skill **already instructs the agent to write as a script**.

## Candidate description

To get per-row eval results (which `evaluation_get` does not return), the skill runs a deterministic flow:

1. Construct the Azure AI Projects SDK client (`get_openai_client()`).
2. `runs.retrieve(...)` for status, then `output_items.list(...)` (SDK auto-paginates).
3. Apply the documented `custom_score` / `{evaluator_name}` dual-entry merge.
4. Persist to `.foundry/results///.json` (`json.dump` with `default=str`).
5. Print passed / failed / errored counts.

This is a strong script candidate because it is:
- **Explicitly meant to be a script** — the skill says "Write a Python script (save to `scripts/`)" and embeds the full implementation, a strong signal it should be a committed, tested script rather than regenerated inline each run.
- **Deterministic and output-heavy** — fixed client construction, fixed pagination, fixed merge, fixed output path, fixed summary; only a few fields of large per-row payloads matter.
- **Multi-caller** — the download is a required step referenced from `evaluate-step.md` (L18–L19) and compared against in `continuous-eval.md`.

**Sketch — `download-eval-results.{sh,ps1}` (or committed Python helper):**
- **Input:** `--eval-id`, `--run-id`, `--environment`.
- **Output:** the persisted JSON path and a passed/failed/errored summary line.

> Failure **clustering** and per-row category review (Steps 4–5 of the same file) are pure judgment and stay in prose, as does the knowledge-cutoff false-negative caveat. Only the download/parse/persist/summary mechanics — the part already written as code — are the candidate.

## Affected file and lines

- [`foundry-agent/observe/references/analyze-results.md` — download/parse/persist/summary (L5–L100)](https://github.com/microsoft/GitHub-Copilot-for-Azure/blob/3890cbfb65c548ce8daa96cabd1d8de63f7bbcca/plugin/skills/microsoft-foundry/foundry-agent/observe/references/analyze-results.md#L5-L100)
- [`foundry-agent/observe/references/evaluate-step.md` — mandates the per-row download (L18–L19)](https://github.com/microsoft/GitHub-Copilot-for-Azure/blob/3890cbfb65c548ce8daa96cabd1d8de63f7bbcca/plugin/skills/microsoft-foundry/foundry-agent/observe/references/evaluate-step.md#L18-L19)

## Next steps

1. **Evaluate the candidate** — confirm the steps are stable and parameterizable, and that the script captures everything the skill needs.
2. **Create both a bash _and_ a PowerShell version** of the script so the skill works across platforms.
3. **Run integration tests** to verify the scripts behave correctly and the skill still completes end-to-end.

## Background Information

### Why replace regular steps with scripts

Replacing a regular, well-defined series of steps with a script can:

- **Reduce token usage** — the skill no longer needs to spell out each command and parse large command output inline; the agent invokes one script and reads a compact result.
- **Improve reliability** — the logic is written and tested once, instead of being re-derived by the agent on every run.
- **Improve determinism** — the same inputs always produce the same steps and output, removing run-to-run variation.
- **Improve speed of execution** — a single script call replaces multiple round-trips of command generation, execution, and large-output parsing.

### Authoring notes for the scripts

- **Reference scripts with markdown links**, not just a bare path to the script file.
- **Include examples** in the skill showing how to run each script (sample invocation with arguments).
- **Briefly explain what each script does** where it is referenced.
- **The script output should explain what it did**, so the agent and user can understand the result without re-inspecting raw command output.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.