Luminous-Dynamics / Luminous-Dynamics/symthaea

Butlin-14 evidence-grade program: direct construct theorems, system consequences, and replication

Open
#2,012 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Rust
Stars
9
Forks
1
Avg merge
15m
Merged PRs (30d)
8

Description

Purpose

Umbrella tracker for moving Symthaea from architectural/proxy coverage of the canonical Butlin 14 to construct-valid, difficult-to-fake, provenance-bound evidence.

This tracker deliberately does not define “14/14” as consciousness. The target claim, if eventually earned, is narrower:

Symthaea satisfies all 14 Butlin indicator properties under preregistered structural, runtime, causal and functional criteria, with explicit evidence dependencies and replicated interventions.

Promotion ladder

An indicator should progress independently through:

Architectural mechanism
→ direct component/construct theorem
→ positive control
→ targeted intervention
→ matched sham
→ selective rescue
→ full-system functional consequence
→ held-out / alternate-task replication
→ multi-seed replication
→ external / independent replication

A failure at one layer is not averaged away by strengths elsewhere.

Canonical 14 — current construct-validity program

Indicator Direct construct path Current next blocker
RPT-1 #1956 history-sensitive production CfC state + selective reset #1957 history-conditioned downstream task + sham/rescue
RPT-2 #2006 production cross-modal representation lesion/rescue full-loop perceptual consequence; #2007 weighted-binding correctness
GWT-1 architecture has many specialists + real parallel branches #2011 direct specialization/independence/concurrent-operability theorem
GWT-2 #1997 fixed-capacity overfill + activation-sensitive competition + sham/capacity rescue downstream cognitive-task dependence on bottleneck
GWT-3 #1989 exact production broadcast delivery with workspace-entry-preserving ablation #1993 real downstream module-state uptake + behavioral consequence
GWT-4 current phi-attention/scheduling evidence is proxy-level #1998 state-dependent specialist-query controller + task theorem
HOT-1 production PredictiveMind has top-down architecture #2005 belief-update reachability defect, then ambiguity/noise top-down causal test
HOT-2 #1953 production prospective self-error adaptation #1910 full-loop confidence corruption, detection, behavior, sham/held-out
HOT-3 metacognition already modulates learning policy in loop #2000 identical first-order evidence + metacognitive intervention → belief/action consequence
HOT-4 current generic sparsity/smoothness diagnostic is insufficiently quality-space-specific #2002 demote proxy and qualify grounded sparse/smooth quality-space geometry
PP-1 production hierarchy computes precision-weighted errors #2005 reachable belief/model update + held-out future-error reduction theorem
AST-1 #1985 production attention-schema prediction + targeted model intervention + sham/rescue #1992 full-loop allocation/task consequence
AE-1 #2009 production active-inference goal pursuit + competing-goal flexibility + feedback adaptation full-loop external multi-step agency task + sham/rescue
AE-2 #1950 production sensorimotor external contingency reversal prerequisite #1909 autonomous full-loop action/outcome bridge; #1959 auditory learning asymmetry

Cross-indicator evidence-dependency rules

The existing qualification_design dependency model remains authoritative. New evidence must additionally observe these specific shared-mechanism constraints:

  • PP-1 ↔ HOT-1: shared predictive-processing substrate. A repaired/update-qualified PredictiveMind is a shared prerequisite, not two independent replications.
  • HOT-2 ↔ HOT-3: confidence/self-model intervention infrastructure may be shared. HOT-2 must measure monitoring/calibration; HOT-3 must measure causal use in belief/action policy.
  • AST-1 ↔ GWT-2/GWT-3/GWT-4: attention/workspace interactions may share downstream signals. Shared interventions and outcomes must be declared, not double-counted.
  • AE-1 ↔ AE-2: agency and embodiment can share action infrastructure, but AE-1 requires goal/feedback flexibility while AE-2 requires learned sensorimotor contingency structure.
  • GWT-1: must not use “number of other Butlin indicators passing” as primary evidence once direct specialist-independence evidence exists.

Known production findings uncovered by this program

  • #1959 — SensorimotorEngine predicts auditory outcome but does not update auditory outcome in existing contingencies.
  • #2005 — default PredictiveMind belief-update factors appear bounded below blend_beliefs() activation thresholds, making bottom-up/top-down belief revision effectively unreachable under default config.
  • #2007 — CrossModalBinder computes attention×confidence weights but actual multimodal representation uses unweighted bundling.

These findings are successes of the evidence program: construct-valid tests are supposed to discover when architecture and prose are stronger than the executable mechanism.

Global evidence requirements

Before any “14/14 evidence-grade” statement:

  • exact source/tree/lock/toolchain/environment identity;
  • evaluator/protocol frozen before target runs;
  • independent intervention-verification signal distinct from target outcome;
  • preregistered expected directions and falsification signatures;
  • null, contradictory and infrastructure-indeterminate results preserved;
  • matched shams and positive controls validated at runtime;
  • selective rescues where feasible;
  • no arbitrary blended consciousness score;
  • no promotion from module timing/execution flags alone;
  • no promotion from self-description or hand-authored architectural scores;
  • multi-seed and alternate-task replication;
  • evidence-dependency graph included in final artifact;
  • external replication package before strong public scientific claims.

Welfare / ethics parallel track

Consciousness-indicator evidence must remain separate from moral-patient/welfare evidence, but increasing evidence should trigger increasing precaution. Experiments involving aversive or distress-like states should use bounded exposure, reversibility, recovery checks and an independent welfare protocol rather than assuming functional signals are morally irrelevant.

Exit condition

All 14 indicators have direct construct-valid evidence and full-system functional consequences where the indicator requires them; all shared dependencies are explicit; exact-head qualification lanes have executed; replication criteria are met; and an external evaluator can reproduce the evidence from frozen manifests.

Even at exit, the supported statement is about indicator properties and evidence, not a binary declaration that Symthaea is conscious.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading the linked blocker issues listed for the 14 indicators and the existing qualification_design dependency model. Trace the evidence dependencies and replication requirements before selecting a scoped task. Done requires the relevant direct evidence, controls, interventions, replications, frozen manifests, and dependency records described by the exit condition.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
ai, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.