microsoft / microsoft/GitHub-Copilot-for-Azure

Replace microsoft-foundry skill eval.yaml parsing (Step 4) with a script

Open
#2,475 1 comment 0 reactions 0 assignees View on GitHub
microsoft-foundry skills untriaged
Dominant language
Python
Stars
250
Forks
204
Avg merge
1d 12h
Merged PRs (30d)
67

Description

## Summary

Copilot has identified a portion of a skill that is a good candidate for replacement with a script.

The candidate is the **eval.yaml parsing (Step 4)** in the `microsoft-foundry` skill (`SKILL.md`, Common Project Context Resolution).

## Candidate description

**eval.yaml parsing (Step 4)**

This step parses `eval.yaml` from the selected agent root and extracts a fixed field set (`agent.name`, `dataset_file`, `evaluators[]`, `name`, `options.eval_model`, `options.pass_threshold`, `max_samples`, `trace_days`, `generation_instruction`). This is deterministic YAML extraction with small output — a script could emit just those fields. Note the *verification* steps (for example confirming evaluators with `evaluator_catalog_get`) remain agent/MCP work, so a script covers the extraction portion only, not the full step.

## Affected file and lines

- [`SKILL.md` — Step 4: Resolve eval.yaml Local Evaluation Intent (L160–L170)](https://github.com/microsoft/GitHub-Copilot-for-Azure/blob/164e0cda7b9d75d6e0d64b235d17bd7ff5e2909e/plugin/skills/microsoft-foundry/SKILL.md#L160-L170)

## Next steps

1. **Evaluate the candidate** — confirm the steps are stable and parameterizable, and that the script captures everything the skill needs.
2. **Create both a bash _and_ a PowerShell version** of the script so the skill works across platforms.
3. **Run integration tests** to verify the scripts behave correctly and the skill still completes end-to-end.

## Background Information

### Why replace regular steps with scripts

Replacing a regular, well-defined series of steps with a script can:

- **Reduce token usage** — the skill no longer needs to spell out each command and parse large command output inline; the agent invokes one script and reads a compact result.
- **Improve reliability** — the logic is written and tested once, instead of being re-derived by the agent on every run.
- **Improve determinism** — the same inputs always produce the same steps and output, removing run-to-run variation.
- **Improve speed of execution** — a single script call replaces multiple round-trips of command generation, execution, and large-output parsing.

### Authoring notes for the scripts

- **Reference scripts with markdown links**, not just a bare path to the script file.
- **Include examples** in the skill showing how to run each script (sample invocation with arguments).
- **Briefly explain what each script does** where it is referenced.
- **The script output should explain what it did**, so the agent and user can understand the result without re-inspecting raw command output.

Contributor guide

Open the contributing guide

Research direction

Start by reading SKILL.md Step 4, Resolve eval.yaml Local Evaluation Intent, at L160–L170, and identify the stable fields the extraction must emit. Create bash and PowerShell scripts that extract those fields, then update the skill to link to them with examples and explanations. Run integration tests to verify cross-platform behavior and that the skill still completes end-to-end.

Written by the indexing model from the issue text.

Assessment

Tech stack
bash, powershell
Domain
tooling
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.