EIDOS episodic reasoning: prediction-outcome tracking for agent evolution
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Stale
- Tech stack
- rust
- Domain
- ai, backend-api-design
Research direction
Start by reading the terraphim_agent_evolution and terraphim_types crates, then review dependencies #597, #599, and #600. The affected areas also include terraphim_hooks and terraphim_mcp_server. Done means predictions and observed outcomes are modeled, stored and matched, accuracy is exposed, and confidence feedback is connected across the stated sources.
Written by the indexing model from the issue text.
Description
Summary
Implement episodic reasoning in terraphim_agent_evolution: every advisory/decision carries a predicted outcome. When the actual outcome is observed, prediction accuracy feeds back into future confidence scoring.
Motivation
Inspired by vibeship-spark-intelligence EIDOS (Episode-based Distillation) loop. terraphim_agent_evolution exists as a crate but is incomplete -- it's supposed to "track agent performance metrics, evolve agent capabilities over time." EIDOS provides the concrete mechanism.
EIDOS Loop
1. PREDICT: "This advisory/action will produce outcome X"
|
2. ACT: Agent or developer takes action
|
3. OBSERVE: Actual outcome Y recorded
|
4. EVALUATE: Compare prediction X vs outcome Y, compute accuracy
|
5. DISTILL: Update confidence for this class of advisory/action
|
(loop)
Concrete Application in Terraphim
Learning Capture Predictions
When terraphim-agent learn surfaces a past learning as advisory:
- Predict: "Following this learning will prevent error type Z"
- Observe: Did the developer follow the advice? Did the error recur?
- Evaluate: If followed and error didn't recur -> prediction validated
- Distill: Increase learning's reliability score
Judge Verdict Predictions
When the judge system issues a verdict:
- Predict: "This code change will cause issue type Z if merged"
- Observe: Was the verdict followed? Did the predicted issue manifest?
- Evaluate: If ignored and issue manifested -> judge was right but ignored
- Distill: Increase judge finding weight for this pattern
Hook Replacement Predictions
When terraphim_hooks replaces text (e.g., npm -> bun):
- Predict: "This replacement will succeed without side effects"
- Observe: Did the subsequent command succeed?
- Evaluate: If replacement caused a failure -> hook rule needs refinement
- Distill: Update hook confidence or add exception rule
Implementation Plan
1. Prediction type in terraphim_types
pub struct Prediction {
pub id: Ulid,
pub source: PredictionSource, // Learning, Judge, Hook, Agent
pub predicted_outcome: String,
pub confidence: f64, // 0.0 - 1.0
pub created_at: jiff::Timestamp,
pub observed_outcome: Option<ObservedOutcome>,
pub accuracy: Option<f64>,
}
2. Prediction tracking in terraphim_agent_evolution
- Store predictions in event store (#597)
- Match outcomes to predictions via correlation IDs
- Compute rolling accuracy per prediction source
- Expose accuracy metrics via MCP resource
3. Confidence feedback
- When prediction accuracy is high (>0.8 over 10+ predictions): increase source confidence
- When prediction accuracy is low (<0.5 over 10+ predictions): decrease source confidence or flag for review
- Confidence scores feed back into advisory ranking and judge dimensional scoring (#600)
Affected Crates
terraphim_agent_evolution(primary -- implement EIDOS loop)terraphim_types(add Prediction, ObservedOutcome types)terraphim_hooks(emit predictions for text replacements)terraphim_mcp_server(expose prediction accuracy metrics)
Dependencies
- #597 Event sourcing (predictions stored as events)
- #599 Enhanced learning capture (learning advisories generate predictions)
- #600 Dimensional verdict scoring (judge verdicts generate predictions)
Estimated Effort
~1 day for prediction types + basic tracking. Feedback loop refinement is ongoing.
Key Insight
Agent evolution is not about self-modifying code (Ouroboros pattern). It's about self-improving knowledge quality. The evolution happens in the advisory confidence scores, not in the source code. This is the right approach for terraphim's deterministic-first philosophy -- the Aho-Corasick engine stays constant, the confidence weights evolve.
- Dominant language
- Rust
- Stars
- 62
- Forks
- 5
- Avg merge
- 2h 27m
- Merged PRs (30d)
- 1
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from terraphim/terraphim-ai
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
terraphim/terraphim-ai#885 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 55/100
terraphim/terraphim-ai#871 ·
-
enhancement
Difficulty 5/5 Over a week Newbie friendliness 25/100
terraphim/terraphim-ai#810 · 2 comments ·
-
enhancement
Difficulty 5/5 Over a week Newbie friendliness 35/100
terraphim/terraphim-ai#729 ·
-
enhancement
Difficulty 5/5 Over a week Newbie friendliness 35/100
terraphim/terraphim-ai#728 ·
All issues in terraphim/terraphim-ai
Similar issues
-
risk:low runtime status:in-progress type:test
Difficulty 1/5 Under an hour Newbie friendliness 92/100
zeroclaw-labs/zeroclaw#11023 ·
-
good first issue refactor
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
EricSpencer00/Resilient#4835 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
bisq-network/bisq-musig#204 ·
-
agent:ready documentation
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
cesarferreira/stax#890 ·