EIDOS episodic reasoning: prediction-outcome tracking for agent evolution

Open
#601 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Stale
Tech stack
rust

Research direction

Start by reading the terraphim_agent_evolution and terraphim_types crates, then review dependencies #597, #599, and #600. The affected areas also include terraphim_hooks and terraphim_mcp_server. Done means predictions and observed outcomes are modeled, stored and matched, accuracy is exposed, and confidence feedback is connected across the stated sources.

Written by the indexing model from the issue text.

Description

architecture enhancement multi-agent

Summary

Implement episodic reasoning in terraphim_agent_evolution: every advisory/decision carries a predicted outcome. When the actual outcome is observed, prediction accuracy feeds back into future confidence scoring.

Motivation

Inspired by vibeship-spark-intelligence EIDOS (Episode-based Distillation) loop. terraphim_agent_evolution exists as a crate but is incomplete -- it's supposed to "track agent performance metrics, evolve agent capabilities over time." EIDOS provides the concrete mechanism.

EIDOS Loop

1. PREDICT: "This advisory/action will produce outcome X"
   |
2. ACT: Agent or developer takes action
   |
3. OBSERVE: Actual outcome Y recorded
   |
4. EVALUATE: Compare prediction X vs outcome Y, compute accuracy
   |
5. DISTILL: Update confidence for this class of advisory/action
   |
   (loop)

Concrete Application in Terraphim

Learning Capture Predictions

When terraphim-agent learn surfaces a past learning as advisory:

  • Predict: "Following this learning will prevent error type Z"
  • Observe: Did the developer follow the advice? Did the error recur?
  • Evaluate: If followed and error didn't recur -> prediction validated
  • Distill: Increase learning's reliability score
Judge Verdict Predictions

When the judge system issues a verdict:

  • Predict: "This code change will cause issue type Z if merged"
  • Observe: Was the verdict followed? Did the predicted issue manifest?
  • Evaluate: If ignored and issue manifested -> judge was right but ignored
  • Distill: Increase judge finding weight for this pattern
Hook Replacement Predictions

When terraphim_hooks replaces text (e.g., npm -> bun):

  • Predict: "This replacement will succeed without side effects"
  • Observe: Did the subsequent command succeed?
  • Evaluate: If replacement caused a failure -> hook rule needs refinement
  • Distill: Update hook confidence or add exception rule

Implementation Plan

1. Prediction type in terraphim_types
pub struct Prediction {
    pub id: Ulid,
    pub source: PredictionSource,     // Learning, Judge, Hook, Agent
    pub predicted_outcome: String,
    pub confidence: f64,              // 0.0 - 1.0
    pub created_at: jiff::Timestamp,
    pub observed_outcome: Option<ObservedOutcome>,
    pub accuracy: Option<f64>,
}
2. Prediction tracking in terraphim_agent_evolution
  • Store predictions in event store (#597)
  • Match outcomes to predictions via correlation IDs
  • Compute rolling accuracy per prediction source
  • Expose accuracy metrics via MCP resource
3. Confidence feedback
  • When prediction accuracy is high (>0.8 over 10+ predictions): increase source confidence
  • When prediction accuracy is low (<0.5 over 10+ predictions): decrease source confidence or flag for review
  • Confidence scores feed back into advisory ranking and judge dimensional scoring (#600)

Affected Crates

  • terraphim_agent_evolution (primary -- implement EIDOS loop)
  • terraphim_types (add Prediction, ObservedOutcome types)
  • terraphim_hooks (emit predictions for text replacements)
  • terraphim_mcp_server (expose prediction accuracy metrics)

Dependencies

  • #597 Event sourcing (predictions stored as events)
  • #599 Enhanced learning capture (learning advisories generate predictions)
  • #600 Dimensional verdict scoring (judge verdicts generate predictions)

Estimated Effort

~1 day for prediction types + basic tracking. Feedback loop refinement is ongoing.

Key Insight

Agent evolution is not about self-modifying code (Ouroboros pattern). It's about self-improving knowledge quality. The evolution happens in the advisory confidence scores, not in the source code. This is the right approach for terraphim's deterministic-first philosophy -- the Aho-Corasick engine stays constant, the confidence weights evolve.

Dominant language
Rust
Stars
62
Forks
5
Avg merge
2h 27m
Merged PRs (30d)
1

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from terraphim/terraphim-ai

All issues in terraphim/terraphim-ai

Similar issues

More Rust issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.