github / github/spec-kit

[Bug]: `/speckit.converge` can declare convergence without auditable acceptance-scenario coverage

未关闭
#3,752 4 条评论 0 个 reaction 已指派 1 人 已被 @BenBtg 认领 在 GitHub 查看
bug-fix fix-proposed severity-high
主要语言
Python
星标
137k
派生
12.3k
平均合并
2 天 12 小时
30 天内合并 PR
159

描述

### Bug Description

`/speckit.converge` already requires one stable intent-inventory key per `FR-###`, buildable `SC-###`, and user-story acceptance scenario such as `US1/AC2`. However, its output contract exposes only a single undenominated line, `Requirements / acceptance criteria checked`.

An agent can therefore assess user stories in bulk, report a plausible aggregate (for example, `16 FR + 12 SC + 4 user stories`), and still emit `converged` without exposing that only 4 user-story rows were checked instead of the 16 acceptance scenarios required by the inventory rule.

This is an observability/closure-contract defect rather than a missing inventory instruction: the prompt tells the agent what to assess, but the report cannot demonstrate whether it did so.

### Steps to Reproduce

1. Run `/speckit.converge` after an implementation pass on a feature whose `spec.md` has multiple user-story acceptance scenarios.
2. Have the assessment aggregate user stories instead of evaluating each scenario key.
3. Observe that the report can satisfy the current summary-metrics text while omitting the number of scenarios assessed.
4. In the observed case, a second pass changed only the granularity from user stories to individual acceptance scenarios and surfaced a missing HIGH finding at `US3/AC4`; a third pass at the same granularity found no remaining gaps.

### Expected Behavior

The convergence report should make completion independently auditable:

- report `checked/total` separately for buildable FRs, buildable SCs, user-story acceptance scenarios, plan decisions, and applicable Constitution MUST principles;
- derive each denominator from the artifacts loaded during the run, not from model memory;
- expose the stable inventory keys with a compact checked/gap/unassessed status;
- allow `converged` only when every applicable category is fully assessed and no findings remain;
- fail closed with a distinct `incomplete_assessment` result when any inventory key is unassessed, without appending a partial Convergence phase to `tasks.md`.

### Actual Behavior

The current template requires the fine-grained inventory internally but only asks for aggregate summary labels. As a result, a report can look complete while concealing a mismatch such as 4 assessed user stories versus 16 required acceptance scenarios. A false `converged` would end the implementation/convergence loop with an unassessed HIGH requirement.

### Specify CLI Version

0.14.3.dev0 (`main` at time of report)

### AI Agent

Not applicable — this is a core command-template contract affecting all agent integrations.

### Operating System

Linux 5.15.147-sun60iw2

### Python Version

Python 3.13.5

### Error Logs

```text
No process error: the defect is a false-completeness reporting path.
```

### Additional Context

Proposed focused change:

1. Update `templates/commands/converge.md` with denominated per-category coverage, an auditable key ledger, and the fail-closed outcome precedence `incomplete_assessment` → `tasks_appended` → `converged`.
2. Document the third outcome in the Agentic SDD reference.
3. Add a deterministic template-contract regression test plus manual command validation.

Searches of existing upstream Issues/PRs found no duplicate for `converge` coverage denominators or incomplete assessment.

**AI disclosure:** This report was prepared by Hermes Agent (OpenAI gpt-5.6-terra) on behalf of @cadugevaerd, using Carlos's observed three-pass evidence and current upstream source inspection.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。