microsoft / microsoft/SkillOpt
section_present rejects numbered or annotated headings, and the optimizer 'fixes' it by forbidding the skill's own heading format
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 17.3k
- Forks
- 1.6k
- Avg merge
- 2d 7h
- Merged PRs (30d)
- 17
Description
Problem
judges._section_present anchors \s*$ immediately after the section name:
pat = re.compile(r"(?im)^\s{0,3}(#{1,6}\s*.*%s|\*\*.*%s.*\*\*\s*:?)\s*$" % (name, name))
So a heading passes only when the name is the last thing on the line:
1 ## Key Risks
1 **Key Risks:**
0 ### 1. Key Risks (Риски) — обзор ← numbered + translated + subtitle
0 ### Key Risks (Risks)
Numbered headings, bilingual headings, and Heading — subtitle are all common in real skill documents, and none of them survive.
Why this is more than cosmetic
My skill's own body demonstrates headings in exactly the rejected style, in its "structure" section:
### 1. Пратигья (Pratijñā) — Тезис
### 2. Хету (Hetu) — Причина
The model faithfully reproduced that format, so all five section_present checks failed on an otherwise correct answer (soft 0.55 with every content check passing). The optimizer then read those failures and proposed:
OVERRIDE: You MUST use EXACTLY these section titles […]. Do NOT append numbers, Latin transliterations, or descriptive text to the headings (e.g. output
### Пратигья, NOT### 1. Пратигья (Pratijñā) — Тезис).
The gate accepted it: 0.682 → 0.852. So a judge artifact produced a rule that forbids the format the skill itself teaches, and it would have been written into the skill permanently had I not read the per-task diff. As a side effect the same rule made the model drop a required citation from another task, which the mean hid (filed separately as the no-regression issue).
This is a concrete instance of the Goodhart pattern discussed in #154, arising purely from a strictness mismatch in one operator.
Possible directions
I did not send a patch because any change here alters scoring for existing task sets, so it seems like a maintainer call:
- Relax the anchor — allow trailing text after the name (
^#{1,6}[^\n]*<name>). Most faithful to what "section present" means, but existing sets that relied on the strict form would start passing more. - Add a
section_containsop and leavesection_presentuntouched. No behaviour change; authors opt in. - Document the strictness in the operator list, so authors know a numbered heading will not match.
My own workaround was replacing every section_present with (?im)^\s{0,3}(?:#{1,6}|\*\*)[^\n]*<name>, which behaves as I expected the built-in to.
Repro
from skillopt_sleep.judges import score_rule_judge
j = {"kind": "rule", "checks": [{"op": "section_present", "arg": "Key Risks"}]}
score_rule_judge(j, "## Key Risks")[0] # 1.0
score_rule_judge(j, "### 1. Key Risks (Риски) — обзор")[0] # 0.0
main @ fdeebaf, Python 3.11.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start at judges._section_present and reproduce the mismatch through score_rule_judge with the two headings shown in the issue. Review existing judge tests and task sets to determine the intended heading semantics and scoring compatibility; done means the chosen behavior is covered by regression tests without breaking existing strict-form checks.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100