microsoft / microsoft/GitHub-Copilot-for-Azure
Define and document the core-plugin evaluation model and quality metrics
- Dominant language
- Python
- Stars
- 250
- Forks
- 204
- Avg merge
- 1d 12h
- Merged PRs (30d)
- 67
Description
## Context
> For the core plugin/skills we own. We will continuously invest in refining the evaluation metrics to answer the following questions:
> - What user prompts do we expect to trigger the skill
> - What do we expect the LLM to do with the skill
> - With the skill, does the LLM accomplish the goal meeting the expectations
> - Without the skill, does the LLM accomplish the goal meeting the expectations
## Scope
- Define the evaluation model covering the four questions above.
- Document quality metrics used to measure core-plugin skill quality.
## Acceptance criteria
- A documented, repeatable evaluation model + metrics exists for the core plugin.
Contributor guide
Research direction
No files, tests, or entry points are named. Start by reviewing the core plugin/skills context and map the four stated questions into a repeatable evaluation model with quality metrics. Done means the model and metrics are documented and can be applied consistently to core-plugin skills.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation, testing-qa
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 38/100