microsoft / microsoft/GitHub-Copilot-for-Azure
[Task]: run comparison workflow on old skills to determine their value
- Dominant language
- Python
- Stars
- 250
- Forks
- 204
- Avg merge
- 1d 12h
- Merged PRs (30d)
- 67
Description
## Skills
Skills to run comparison tests for:
- [x] airunway-aks-setup
- [x] appinsights-instrumentation
- [ ] azure-ai (low priority, a parnter team is doing this for us)
- [ ] azure-aigateway (low priority, no e2e tests for comparison exists)
- [ ] azure-cloud-migrate
- [ ] azure-compliance (low priority, no e2e tests for comparison exists)
- [ ] azure-compute
- [ ] azure-cost
- [ ] azure-diagnostics
- [ ] azure-enterprise-infra-planner
- [ ] azure-kubernetes
- [ ] azure-kusto
- [ ] azure-messaging
- [ ] azure-quotas
- [ ] azure-reliability
- [x] azure-resource-lookup
- [ ] azure-resource-visualizer
- [ ] azure-storage
- [ ] azure-upgrade
- [ ] entra-agent-id
- [ ] entra-app-registration
- [ ] python-appservice-deploy
These skills have been recently visited or are considered critical so they are exempt from the evaluation:
- azure-app-onboard
- azure-app-onboard-prereq
- azure-deploy
- azure-prepare
- microsoft-foundry
## Criteria
We want to evaluate the value of the skills from 3 different perspectives:
- quality: does the skill make the agent generate more useful content
- consistency: does the skill influence the agent to have a more consistent behavior pattern
- cost: does the skill reduce token usage or increase token usage, by how much
Based on the test run trajectories, answer the following questions:
1. For trajectories without the skill, does the agent do an equivalently good job compared to the trajectories with the skill for the same model?
2. For trajectories with the skill, does the agent have a consistent behavior pattern? Does the agent have a consistent behavior pattern without the skill?
3. How much additional token does the agent use with the skill?
Contributor guide
Assessment
This issue has not been assessed yet.