Support temperature/sampling parameters in plugin SKILL.md metadata
- Vorherrschende Sprache
- Shell
- Sterne
- 11.2k
- Forks
- 1.9k
- Ø Merge
- 14 Std. 16 Min.
- Gemergte PRs (30 T.)
- 6
Beschreibung
### Describe the feature or problem you'd like to solve
Entire teams writing official tech docs but can't change the temperature or top-p
### Proposed solution
Plugin authors building documentation-authoring skills need low-temperature inference to minimize fabrication and hallucination. Currently, there is **no way** for a plugin or skill to specify `temperature`, `top_p`, or other sampling parameters -- not in `plugin.json`, `agency.json`, or SKILL.md frontmatter metadata.
The only workaround is embedding prompt-level instructions like "behave as if temperature is 0.0-0.2," which is unreliable because the model may not honor behavioral temperature constraints the same way it honors API-level parameters.
Allow SKILL.md frontmatter `metadata` to include sampling parameters that Copilot CLI passes through to the inference API:
```yaml
---
name: write-concept
description: Author a concept article with verification
metadata:
temperature: 0.1
top_p: 0.9
max_tokens: 8192
---
```
Or alternatively, support this at the plugin level in `plugin.json`:
```json
{
"name": "author-pro",
"modelParameters": {
"temperature": 0.1,
"top_p": 0.9
}
}
```
Either approach would let plugin authors tune inference behavior for their use case -- low temperature for factual documentation, higher temperature for creative brainstorming, etc.
1. **Prompt-level behavioral constraints** -- "Respond as if temperature is set to 0.1." Works partially but is not equivalent to API-level control. Models still exhibit higher variance than a true low-temperature API call.
2. **MCP server with sampling** -- Build an MCP server that uses the `sampling/createMessage` protocol method with temperature parameters. This is architecturally heavier and adds latency, but could work if sampling parameters are passed through to the inference backend.
3. **Direct API calls via MCP** -- Build an MCP server that calls Azure OpenAI directly with explicit temperature. This works but requires separate API credentials, deployment management, and cost -- defeating the purpose of an integrated plugin system.
### Example prompts or workflows
- This affects any plugin where output accuracy matters more than creativity (documentation, compliance, security review, technical writing).
- The `task` tool already supports a `model` parameter for sub-agents. Extending this pattern to include `temperature` would be consistent.
- MCP sampling events (`sampling.requested` / `sampling.completed`) already exist in the session events schema, suggesting the infrastructure may partially exist.
### Additional context
Beitragsleitfaden
Rechercherichtung
Beginne damit nachzuverfolgen, wie die Frontmatter-Metadaten von SKILL.md, plugin.json und agency.json geparst werden, und vergleiche diesen Pfad anschließend mit dem bestehenden Modellparameter des Task-Tools. Überprüfe das Schema der Sitzungsereignisse auf die genannten MCP-Sampling-Ereignisse und ermittle, wo Sampling-Parameter die Inference API erreichen können. Als abgeschlossen gilt, wenn ein unterstützter Metadatenpfad definiert ist und temperature, top_p und max_tokens konsistent durchgereicht werden.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- github, shell
- Bereich
- ai, api, cli, tooling
- Issue-Typ
- Feature
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Ruhig
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 48/100