[Dream Cycle 2026-09-12] intelligence: findSimilar() conflates learned confidence with query-similarity, inverting distillLearning()'s reliability gate + capabilities,memory scan
- Dominant language
- TypeScript
- Stars
- 72.7k
- Forks
- 8.6k
- Avg merge
- 2d 23h
- Merged PRs (30d)
- 83
Description
## 1. Tonight's Rotation
```
DATE=2026-09-12
DEEP=intelligence
SCAN=capabilities,memory
SLOT=2 (DAYINT % 5 == 2)
COMMIT=39e0b0540c9b018174955fc8a21f355bbac26c6a
No bonus deep-dive triggers (DAYINT%25=12, %75=37)
```
## 2. Ledger Check
Inspected last 14+ rows of `docs/dream-cycle/LEDGER.md` and verified GitHub state via MCP tools (no `gh` CLI in this environment) for the last 7 nights:
| Date | Surface | Issue/PR | State |
|---|---|---|---|
| 09-01 | security | #3151/#3152 | OPEN/draft, stacked on unmerged #3139 |
| 09-02 | intelligence | #3159/#3160 | OPEN/draft |
| 09-03 | memory | #3168/#3169 | **MERGED** |
| 09-05 | performance | #3183/#3184 | **MERGED** |
| 09-07 | intelligence | #3220/#3221 | **MERGED** |
| 09-08 | memory | #3231/#3232 | OPEN/draft, **mergeable_state: dirty (merge conflict)** — new since 09-09's check |
| 09-09 | swarm | #3242/#3243 | **MERGED** |
Trailing-7-night merge rate: 4/7 (up from the historical "0 of last 14" low). No forced small-patch bias triggered by STEP 1.1, but the general small-scoped-patch discipline was kept anyway. No 3x-duplicate-direction on the `intelligence` surface (Thompson-decay #3049, EWC-gate #3110, LearningBridge #3160, tier-forwarding #3221 are all distinct bugs) — no surface substitution needed.
**Prior gist score (STEP 1.2):** scored the 2026-09-09 gist 9/10 (2 benchmark-evidence, 2 competitor-rows [6 rows], 2 executable-recommendations, 2 valid-witness, 0 for <1500 words [it ran long], 1 for genuine novelty).
## 3. Deep Dive Findings
Five research roles ran in parallel (deep researcher, 2 scan researchers, competitor analyst, independent architecture reviewer — see `docs/dream-cycle/dream-gist-2026-09-12.md` for full detail).
**Selected finding** (architecture reviewer): `LocalReasoningBank.findSimilar()` (`v3/@claude-flow/cli/src/memory/intelligence.ts:604-638`, pre-fix) overwrote every returned pattern's `confidence` (learned reliability) with the per-query cosine similarity score — `return { ...s.pattern, confidence: s.score }`. This broke two things: (a) `distillLearning()`'s "only distill from high-confidence matches" gate (`match.confidence < 0.5`, line 357) tested query-similarity instead of reliability, inverting its stated intent; (b) the already-exported `PatternMatch extends Pattern { similarity: number }` contract (line 1195) that `findSimilarPatterns()` realizes could only ever return `similarity === confidence`, via an unsafe type-cast fallback (`(r as unknown as {similarity?:number}).similarity ?? r.confidence ?? 0.5`).
**Deep researcher's 5 additional scored candidates** (not selected tonight, see gist for full detail and scores): `priorDecay` distribution-shift fix built/tested/benchmarked but never wired to any config/env (score 4.55); MoE router missing an explore/exploit toggle unlike its sibling Q-learning router (4.20); `determineEvolutionType()` mislabels neutral/declining pattern evolution as `'improvement'`, duplicated in 2 files (4.15); `ReasoningBank.distill()` only distills successful/partial trajectories, never failure-derived memory despite the paper it's named after (arXiv:2509.25140) arguing for both (3.95); inconsistent cosine-similarity clamping across ~5 duplicate implementations (3.60). The selected finding scores 4.85 under the same rubric (5/5/5/4/5/5) — highest of all candidates, no override needed.
**Independent re-verification**: confirmed #3110 (EWC-gate wrong Fisher dims) and #3160 (LearningBridge reward-blindness) are still present on `main` exactly as the ledger states — both PRs remain open/draft.
## 4. Hypothesis
> Given `LocalReasoningBank.findSimilar()` returning pattern objects whose `confidence` field is overwritten with the per-query cosine similarity score, when the candidate change stops overwriting `confidence` and instead returns the cosine score as a distinct `similarity` field (satisfying the pre-existing `PatternMatch` contract), then (a) `distillLearning()`'s gate should correctly test learned reliability rather than query-similarity, and (b) `findSimilarPatterns()` should return genuinely distinct `confidence`/`similarity` values, subject to: (1) `memory-bridge.ts`'s search-ranking consumer, which wants query-relevance not reliability, is migrated to prefer the new `similarity` field; (2) `endTrajectory()`'s RL-update consumer (which already re-fetches `pattern.confidence` via `bank.get()`) is unaffected; (3) all existing tests remain green; (4) fully deterministic, $0 evaluation cost.
Frozen before evaluation; not modified after seeing results.
## 5. Evaluation Receipt
**evaluated: accepted.** Real evaluator: Vitest 4.1.8, deterministic, $0, zero LLM calls. New file: `v3/@claude-flow/cli/__tests__/intelligence-confidence-similarity-conflation.test.ts` (4 tests).
Baseline vs. candidate, isolated via `git stash` of the 2 source files (test kept): all 4 tests fail against baseline for the exact predicted mechanism (confidence overwritten by cosine 1.0/~0.45; DISTILL gate off by exactly the extra LoRA bump a wrongly-included pattern receives; public-API field conflation), pass on candidate.
Full `@claude-flow/cli` suite both ways: baseline 91 failed/149 passed/1 skipped (241 files), candidate 90 failed/150 passed/1 skipped — the only failed-file-set difference is the new test file itself; the other 90 are identical pre-existing environmental failures (unbuilt monorepo sibling packages, e.g. `@claude-flow/neural` Vite-resolution failures). `tsc --noEmit`: 463 pre-existing errors, byte-identical count/lines both ways.
## 6. Darwin Results
Skipped — binary correctness/API-contract fix, not a continuous parameter with a fitness gradient for Darwin's real interface (`npx ruvector harness darwin --help`, confirmed available, `@metaharness/darwin@0.9.2`) to search over.
## 7. Flywheel Evidence
No `.claude-flow/flywheel/` state or signed `@metaharness/flywheel` bundle in this repo (confirmed via `npx ruvector harness flywheel --help`). Evidence retained as: 4 new tests, gist, this issue, the stash-isolated comparison, and the independent adversarial critique — consistent with every accepted night since 2026-08-18.
## 8. Reward Hack Check
No standalone reward-hack CLI reachable (`@metaharness/weight-eft` is a LoRA-distillation tool, not a diff scanner). Manual checklist, independently re-verified by the adversarial critic: no test weakened, no gold data touched, no cherry-picking (all 90 pre-existing failures disclosed), no seed manipulation (deterministic hand-crafted embeddings), $0 cost, discriminating power confirmed real (fails on baseline for the stated reason).
## 9. Security Review
Not security-sensitive: in-process ranking/confidence bookkeeping only, no new I/O/network/credential/filesystem surface. Independently confirmed by the adversarial critic. The `memory-bridge.ts` migration's `??` fallback chain was specifically checked against a `||` footgun on a legitimate `0`-similarity value — `??` is correct.
## 10. Scan Findings: capabilities
Ruflo's swarm permission/audit system (`v3/@claude-flow/cli/src/permission/{permission-set,permission-audit}.ts`) writes grants but can only ever emit a `'granted'` audit event (3 of 4 declared event kinds — `checked`/`denied`/`revoked` — are dead code, zero call sites in `v3/`), and nothing reads `.swarm/permissions.jsonl` back to gate any tool call, path, or network host. The module's own doc comment is candid it's "a METADATA + AUDIT layer, not a runtime sandbox," but the audit-log shape implies more enforcement than exists. Comparison: OpenAI Agents SDK's `tool_input_guardrail`/`tool_output_guardrail` decorators run inline in the actual tool-call path (Grade A). Not selected tonight (SCAN surface); flagged for a future `capabilities` night.
## 11. Scan Findings: memory
`@claude-flow/memory`'s same-key upsert path is broken across all three store layers: `generateMemoryId()` never produces a deterministic/reusable ID, so `AgentDBAdapter.store()` never evicts the prior occupant of a `(namespace,key)` pair; `SQLiteBackend`'s schema has no `UNIQUE(namespace,key)` constraint, so repeated writes accumulate duplicate rows; `HybridBackend.getByKey()` is hard-wired to the (unordered) SQLite path and can return a stale copy. A working precedent already exists in the same repo (`memory-initializer.ts`'s real `UNIQUE(namespace,key)` + `removeHNSWEntriesByKey()`) that the newer package never carried forward. External comparison: Mem0's v3 ADD-only pipeline hits the documented same failure shape (Grade A, mem0ai/mem0 issues #4956/#5867/#4896). Not selected tonight; key-collision pairs are noted as a free, ground-truth-labeled corpus for PR #3232's still-open near-dup-threshold tuning gap.
## 12. Competitors Reviewed
LangGraph, Microsoft Agent Framework/AutoGen, CrewAI, OpenAI Agents SDK, Google ADK (mandated floor); Letta, GEPA/DSPy, ACE (specific non-floor comparisons); Mem0, Qdrant/Weaviate/Milvus/LanceDB (memory scan); OpenAI Agents SDK guardrails, LangGraph third-party permission middleware (capabilities scan). Full graded table in the gist — re-verifies "no mainstream framework ships human-gated evolutionary self-improvement" with 2026-dated primary sources; notes the ecosystem arguably regressed (OpenAI sunsetting Agent Builder/Evals 2026-11-30).
## 13. Gist
`docs/dream-cycle/dream-gist-2026-09-12.md` (this branch). No gist-creation tool is reachable in this execution environment (GitHub MCP tools cover issues/PRs/repos, not gists) — committed to `docs/dream-cycle/` instead, matching established repo convention for these reports.
## 14. Witness
| Field | Value |
|---|---|
| Session commit | `39e0b0540c9b018174955fc8a21f355bbac26c6a` |
| Gist SHA-256 (pre-witness content) | `cb5ee8eff4a9ef7d0123ea74e3d6c30dba3bf0c2922dd1db0235a55ffc633a09` |
| Witness stamp | `cb60b43b6c2130b41626657fc19395214d2ee929e0aff4fd4ee05468cd4bca57` |
## 15. Recommendation
Merge the linked draft PR (human review required). Pick up `priorDecay` wiring (candidate 1) for a future `intelligence` night. Prioritize merging the growing backlog of ACCEPT-scoped/ACCEPT drafts (#3110, #3152, #3160, #3232 all still open). Consider a themed `memory` night combining the upsert-bug fix with PR #3232's near-dup corpus gap. Consider a `capabilities` night to wire real enforcement (or explicitly document audit-only status) for the swarm permission system.
Contributor guide
Research direction
Start with v3/@claude-flow/cli/src/memory/intelligence.ts, especially findSimilar(), distillLearning(), and the PatternMatch definition. Review memory-bridge.ts and run v3/@claude-flow/cli/__tests__/intelligence-confidence-similarity-conflation.test.ts alongside the existing CLI tests. Done means confidence remains learned reliability, similarity represents query relevance, and the focused tests pass.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- typescript
- Domain
- ai, backend
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 25/100