dotnet / dotnet/skills

[skills-eval] dotnet-maui: 8 skills, 48% pass — 1 P0, 4 to strengthen, 3 keep

Open
#895 4 comments 0 reactions 3 assignees Claimed by @jfversluis View on GitHub
area-dotnet-maui area-maui skill task Triaged
Dominant language
C#
Stars
5.4k
Forks
415
Avg merge
1d 4h
Merged PRs (30d)
84

Description

> ## 🔄 EDITED 2026-07-19 — correction per #909
>
> **This finding has been revised.** The errored trials that drove the **FIX-RELIABILITY (P0)** flag(s) below were shown in [#909](https://github.com/dotnet/skills/issues/909) to be **judge-side infrastructure failures** — a disabled/throttled PAT (`CAPIError 400 organization disabled`) and Vally's `session.idle` judge timeout — **not** fixture nondeterminism. Verdict scores already exclude errored trials, so re-running the classifier with those judge errors dropped reclassifies the affected skill(s):
>
> | Skill | ~~Was~~ | Now |
> | --- | --- | --- |
> | `maui-shell-navigation` | ~~FIX-RELIABILITY (P0)~~ | **STRENGTHEN** |
>
> Struck-through text below is the **superseded** finding — in particular the "pin fixture SDK/tool versions" remediation does **not** apply. The title's P0 count has been updated. No eval re-run is required to correct the scores; recovering the lost judgments (optional) means **re-judging** those trials, not re-running the skill.

## Context — cross-family skill evaluation: `dotnet-maui`

This issue is **self-contained**: it captures everything a skill author needs to act on the `dotnet-maui` plugin without opening the full report.

**What this measures.** Every runnable skill in `dotnet/skills` was run through [Vally](https://github.com/microsoft/evaluate) (0.7) on a **cross-family matrix**: 5 executor model families — `opus-4.8`, `gpt-5.5`, `sonnet-4.6`, `haiku-4.5`, `mai-flash` — each judged by a **different family** (judge ≠ executor; default judge = latest Opus, or GPT when Opus is the executor). For every skill, a **skilled** run is compared against a **baseline** (no-skill) run and scored per executor. This removes single-model and self-judging bias, so a skill that only helps one model family — or only its own family's judge — is visible.

- **Data source:** cross-family CI grid (run `29228914412` + backfills) — 5 executors × 85 runnable skills, **419 scored cells** over 84 skills.
- **Row grain:** one row per **skill**, aggregated across its (up to 5) executor cells. `avgN` is the mean trial count behind the cells (trials 1–17; higher = more statistically trustworthy). `thin-N` flags directional-only rows.

### `dotnet-maui` at a glance (portfolio scorecard)

| Plugin | Skills | Cells | Pass | Impact | Tie-trials | Err | avg ΔTok | avg ΔTurns | avg ΔTools | Headline |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | --- |
| **dotnet-maui** | 8 | 40 | 48% | 0.454 | 26 | 1 | +7,766 | +0.32 | +0.38 | Strong impact, `dotnet-maui-doctor` exemplar |

**Insights.** 8 skill(s); mean impact **0.45**; 3/8 help ≥1 frontier model. Exemplars (win incl. frontier): `dotnet-maui-doctor`. Weak-model-only wins (frontier misses): `maui-safe-area`. Weakest: `maui-collectionview` (0.12).

~~**Address first:** `maui-shell-navigation` — reliability; `maui-collectionview` — regression.~~

**Address first:** `maui-collectionview` — regression. _(`maui-shell-navigation` reclassified to STRENGTHEN — see EDIT above.)_

### How to read the table

Each skill is scored **skilled vs. baseline** on these axes:

| Signal | Column | What it means | Good |
| --- | --- | --- | --- |
| Breadth | `Families ✓` (n/5) | How many of the **5 model families** the skill helps (per-family pass count, max `5/5`) | 4–5 / 5 |
| Where | `Passed on` | *Which* families passed. **Frontier = latest Opus + latest GPT** are **bold**; Sonnet 4.6 / Haiku 4.5 / MAI Flash are mid/low-weight | frontier ✓ |
| Magnitude | `Impact` (−1…+1) | How strongly the judge prefers *skilled* over *baseline* | ≥ 0.4 |
| Decisiveness | `Ties▵` | Trials where the judge saw **no difference** → skill is *inert* | low |
| Safety | `Loss▵` | Trials where **skilled was WORSE** than baseline → skill misfires | ~0 |
| Reliability | `Err` | Trials that errored/crashed in setup or judging | 0 |
| Efficiency | `ΔTok` / `ΔTurns` / `ΔTools` | Extra tokens / agent turns / tool calls vs baseline | ≤ 0 |
| Confidence | `avgN` | Mean trials behind the verdict; low N = directional only | ≥ 3 |
| Invocation | `Call%` | Share of skilled trials where the model actually **invoked the skill** | ~100% |

> **`Families ✓` is cell-level (max `5/5`); `Ties▵`/`Loss▵` are trial-level tallies summed across *all* families** (including the ones where the skill failed). A high `Families ✓` next to non-zero `Loss▵` is not a contradiction — see `Passed on` and the Action text for where losses landed.

**Action buckets** (each skill has one primary action; `[flags]` note secondary concerns):

| Bucket | Priority | Meaning |
| --- | --- | --- |
| FIX-RELIABILITY | 🔴 P0 | Errored trials / no verdict — stabilize the harness before trusting the score |
| FIX-DISCOVERY | 🔴 P0 | Model doesn't invoke it (`Call% < 50%`) — a triggering/description problem |
| FIX-REGRESSION | 🔴 P0 | Skilled is worse than baseline on many trials — the skill misfires |
| ADD-DECISIVENESS | 🟠 P1 | Called ~100% but ties dominate, ~0 impact — inert; needs sharper behavioral steps |
| TRIM-COST | 🟠 P1 | Passes but with heavy token/turn overhead — trim verbosity |
| EXEMPLAR | 🟢 keep | Broad, strong, reliable win — use as a template |
| EFFICIENT-WIN | 🟢 protect | Wins *and* cuts turns/tools — the ideal shape |
| KEEP-POLISH | 🟢 | Solid majority win; minor polish + more trials |
| STRENGTHEN | 🟡 P2 | Marginal/mixed lift — sharpen triggers & success criteria |

### Per-skill actions

| Skill | Families ✓ | Passed on (frontier **bold**) | Impact | Ties▵ | Loss▵ | Err | avgN | Call% | ΔTok | ΔTurns | ΔTools | Action |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| `dotnet-maui-doctor` | 5/5 | **Opus**, **GPT**, Sonnet, Haiku, MAI | 0.74 | 4 | 1 | 0 | 7.6 | 98% | +46,929 | +1.96 | +3.57 | **EXEMPLAR** · ✅ **Template-worthy.** What's good: broad (passes 5/5, incl. frontier Opus, GPT). Keep as-is; lift its structure (crisp triggers + imperative steps) into weaker siblings. |
| `maui-safe-area` | 2/5 | Sonnet, Haiku | 0.70 | 2 | 2 | 0 | 3.8 | 95% | -40,938 | -2.2 | -3.07 | **EFFICIENT-WIN** · ✅ **Better *and* cheaper.** Wins on Sonnet, Haiku while saving 2.2 turns, 3.07 tools, 40,938 tok. What's good: it adds value without bloat — the ideal shape. Protect the brevity; don't let it grow. ⚠️ **But both frontier models miss** — this may only be helping mid/low-weight models. Treat as a weak-model win, not a universal one; check the frontier trajectories before templating. *([both-frontier-miss], [efficient])* |
| `maui-app-lifecycle` | 3/5 | Sonnet, Haiku, MAI | 0.59 | 4 | 0 | 0 | 4 | 95% | +5,479 | +0.3 | +0.1 | **KEEP-POLISH** · 🟢 **Solid majority win** on Sonnet, Haiku, MAI. Losses/ties are concentrated in the **non-passing** cells (Opus, GPT), not the 3 passing ones — the passes are clean. To firm up: (1) raise trials on Opus, GPT; (2) investigate why a frontier model (Opus, GPT) didn't pass — that's the highest-value gap; (3) minor wording polish only. *([both-frontier-miss])* |
| `maui-shell-navigation` | 2/5 | **GPT**, Haiku | 0.50 | 2 | 1 | 1 | 3.8 | 90% | +15,144 | +0.85 | +1.05 | ~~**FIX-RELIABILITY** · 🔴 **1 errored trial(s)** — verdict can't be trusted until setup is deterministic. This is an infra/harness fix, not a content one: (1) capture the failing trial's stderr; (2) pin tool/SDK versions in the fixture; (3) add a setup smoke-check before scoring. *([reliability], [frontier-miss])*~~

🔄 **EDIT (see #909):** **STRENGTHEN** · 🟡 Marginal (2/5 — GPT, Haiku), mostly ties. The 1 errored trial was a **judge-side disabled-PAT failure** (mai), not fixture nondeterminism — this is a decisiveness/content problem, not a reliability one. Sharpen the trigger; add opinionated steps. *([frontier-miss])* |
| `maui-theming` | 2/5 | Sonnet, Haiku | 0.34 | 5 | 2 | 0 | 3.8 | 100% | +5,278 | +0.25 | +0.2 | **STRENGTHEN** · 🟡 **Marginal/mixed** (passes 2/5 — Sonnet, Haiku). Diagnosis: mostly *ties* — too generic/non-prescriptive. Try: (1) sharpen the trigger so it fires only where it wins; (2) add 1–2 opinionated, concrete steps that change behaviour; (3) add trials to separate signal from noise. If frontier models never benefit, scope it explicitly to weaker models or reconsider its value. *([both-frontier-miss], [frontier-regressed])* |
| `maui-data-binding` | 2/5 | **GPT**, Sonnet | 0.32 | 5 | 1 | 0 | 4 | 95% | +9,699 | +0.25 | +0.3 | **STRENGTHEN** · 🟡 **Marginal/mixed** (passes 2/5 — **GPT**, Sonnet). Diagnosis: mostly *ties* — too generic/non-prescriptive. Try: (1) sharpen the trigger so it fires only where it wins; (2) add 1–2 opinionated, concrete steps that change behaviour; (3) add trials to separate signal from noise. *([frontier-miss], [frontier-regressed])* |
| `maui-dependency-injection` | 2/5 | Sonnet, Haiku | 0.32 | 2 | 4 | 0 | 4 | 95% | +10,609 | +0.6 | +0.35 | **STRENGTHEN** · 🟡 **Marginal/mixed** (passes 2/5 — Sonnet, Haiku). Diagnosis: misses both frontier models — likely assumes context they solve unaided. Try: (1) sharpen the trigger so it fires only where it wins; (2) add 1–2 opinionated, concrete steps that change behaviour; (3) add trials to separate signal from noise. If frontier models never benefit, scope it explicitly to weaker models or reconsider its value. *([both-frontier-miss])* |
| `maui-collectionview` | 1/5 | Haiku | 0.12 | 2 | 6 | 0 | 4 | 85% | +9,928 | +0.55 | +0.55 | **FIX-REGRESSION** · 🔴 **Regresses ~30% of trials** (worse than baseline), losing on Opus, MAI. The skill is over-applying. Fixes: (1) add explicit **stop-conditions** ("do NOT act when…"); (2) narrow the trigger to the exact scenario it helps; (3) demote prescriptive edits to *suggestions the agent can decline*. Prioritise the frontier miss (Opus, GPT). *([both-frontier-miss], [frontier-regressed])* |

---

Generated from the cross-family Call-to-Action report (`CALL-TO-ACTION.md` §4–§5; companion `IMPACT-ANALYSIS.md`). Regenerate the underlying tables with `node deep-metrics.mjs "$env:TEMP\cf-ci" agg-ci` → `node gen-cta-tables.mjs agg-ci`. Numbers are directional where `avgN` is low; treat single-trial cells as hypotheses to confirm with more runs.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.