terraphim / terraphim/terraphim-ai

Epic: ToCS-based agent evaluation framework for ADF

Open
#691 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Rust
Stars
62
Forks
5
Avg merge
2h 27m
Merged PRs (30d)
1

Description

Context

No existing ADF mechanism measures whether agents genuinely understand the codebases they operate on, or merely execute surface-level patterns. Theory of Code Space (ToCS, arXiv:2603.00601) provides a 4-dimension evaluation framework for exactly this.

Proposal

Implement ToCS-inspired evaluation to measure ADF agent effectiveness across four dimensions:

Evaluation Dimensions
Dimension What It Measures Metric
Construct Does the agent build an accurate dependency map? Edge F1 by type (IMPORTS, CALLS_API, REGISTRY_WIRES, DATA_FLOWS_TO)
Revise Does the agent update beliefs when code changes? Belief revision score (delta accuracy after code change)
Exploit Can the agent predict impact of changes? Counterfactual probe accuracy
Constraints Does the agent discover architectural rules? Invariant discovery F1 vs CLAUDE.md/domain model rules
Implementation Phases
  1. Phase 0: Run ToCS benchmark against terraphim-ai workspace with current agents (baseline)
  2. Phase 1: Add periodic cognitive map probing -- every N tool calls, externalise understanding as structured JSON
  3. Phase 2: Compare probes against ground truth (KG-derived dependency graph) to compute scores
  4. Phase 3: Feed scores to NightwatchMonitor as new signal type (alert on degradation)
Cognitive Map Probing
  • Injected via PreToolUse hooks (Agent SDK) or system messages (subprocess)
  • Agent outputs structured JSON: nodes (modules), edges (dependencies, typed), confidence scores
  • Compared against ground truth from terraphim KG + tree-sitter analysis
Key Insight from ToCS Research
  • Aho-Corasick automata cover ~67% of edges (IMPORTS level)
  • CALLS_API (~17%) and DATA_FLOWS_TO (~7%) require semantic understanding
  • Some models show "catastrophic belief collapse" -- losing knowledge between probes
  • Evaluation framework should be built BEFORE KG enrichment (measure first, improve later)
Sub-issues (to be created during design phase)
  • Run ToCS baseline against terraphim-ai workspace
  • Implement cognitive map probe injection and collection
  • Implement belief stability monitoring (successive probe comparison)
  • Integrate evaluation scores with NightwatchMonitor
  • KG enrichment with tree-sitter call graph (after baseline confirms gap)

References

  • ToCS paper: https://arxiv.org/abs/2603.00601
  • ToCS repo: https://github.com/che-shr-cat/tocs
  • KB article: cto-executive-system/knowledge/external/context-engineering/tocs-theory-of-code-space-benchmark.md
  • Expansion plan: cto-executive-system/plans/tocs-terraphim-ai-evaluation-plan.md
  • ADF plan: cto-executive-system/plans/adf-architecture-improvements.md (item 3.1)
  • Depends on: #689 (Agent SDK migration for hook-based probe injection)
  • Related: #682 (Pi eval epic), #687 (steering queues)

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with cto-executive-system/plans/tocs-terraphim-ai-evaluation-plan.md, cto-executive-system/plans/adf-architecture-improvements.md, and the referenced ToCS paper and repository; review dependency #689 before planning hook-based work. Done requires the design to be split into scoped sub-issues covering the baseline, probe collection, scoring, belief monitoring, NightwatchMonitor integration, and later KG enrichment.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
ai, testing
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.