Luminous-Dynamics / Luminous-Dynamics/symthaea
Butlin-14 evidence-grade program: direct construct theorems, system consequences, and replication
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 9
- Forks
- 1
- Avg merge
- 15m
- Merged PRs (30d)
- 8
Description
Purpose
Umbrella tracker for moving Symthaea from architectural/proxy coverage of the canonical Butlin 14 to construct-valid, difficult-to-fake, provenance-bound evidence.
This tracker deliberately does not define “14/14” as consciousness. The target claim, if eventually earned, is narrower:
Symthaea satisfies all 14 Butlin indicator properties under preregistered structural, runtime, causal and functional criteria, with explicit evidence dependencies and replicated interventions.
Promotion ladder
An indicator should progress independently through:
Architectural mechanism
→ direct component/construct theorem
→ positive control
→ targeted intervention
→ matched sham
→ selective rescue
→ full-system functional consequence
→ held-out / alternate-task replication
→ multi-seed replication
→ external / independent replication
A failure at one layer is not averaged away by strengths elsewhere.
Canonical 14 — current construct-validity program
| Indicator | Direct construct path | Current next blocker |
|---|---|---|
| RPT-1 | #1956 history-sensitive production CfC state + selective reset | #1957 history-conditioned downstream task + sham/rescue |
| RPT-2 | #2006 production cross-modal representation lesion/rescue | full-loop perceptual consequence; #2007 weighted-binding correctness |
| GWT-1 | architecture has many specialists + real parallel branches | #2011 direct specialization/independence/concurrent-operability theorem |
| GWT-2 | #1997 fixed-capacity overfill + activation-sensitive competition + sham/capacity rescue | downstream cognitive-task dependence on bottleneck |
| GWT-3 | #1989 exact production broadcast delivery with workspace-entry-preserving ablation | #1993 real downstream module-state uptake + behavioral consequence |
| GWT-4 | current phi-attention/scheduling evidence is proxy-level | #1998 state-dependent specialist-query controller + task theorem |
| HOT-1 | production PredictiveMind has top-down architecture |
#2005 belief-update reachability defect, then ambiguity/noise top-down causal test |
| HOT-2 | #1953 production prospective self-error adaptation | #1910 full-loop confidence corruption, detection, behavior, sham/held-out |
| HOT-3 | metacognition already modulates learning policy in loop | #2000 identical first-order evidence + metacognitive intervention → belief/action consequence |
| HOT-4 | current generic sparsity/smoothness diagnostic is insufficiently quality-space-specific | #2002 demote proxy and qualify grounded sparse/smooth quality-space geometry |
| PP-1 | production hierarchy computes precision-weighted errors | #2005 reachable belief/model update + held-out future-error reduction theorem |
| AST-1 | #1985 production attention-schema prediction + targeted model intervention + sham/rescue | #1992 full-loop allocation/task consequence |
| AE-1 | #2009 production active-inference goal pursuit + competing-goal flexibility + feedback adaptation | full-loop external multi-step agency task + sham/rescue |
| AE-2 | #1950 production sensorimotor external contingency reversal prerequisite | #1909 autonomous full-loop action/outcome bridge; #1959 auditory learning asymmetry |
Cross-indicator evidence-dependency rules
The existing qualification_design dependency model remains authoritative. New evidence must additionally observe these specific shared-mechanism constraints:
- PP-1 ↔ HOT-1: shared predictive-processing substrate. A repaired/update-qualified
PredictiveMindis a shared prerequisite, not two independent replications. - HOT-2 ↔ HOT-3: confidence/self-model intervention infrastructure may be shared. HOT-2 must measure monitoring/calibration; HOT-3 must measure causal use in belief/action policy.
- AST-1 ↔ GWT-2/GWT-3/GWT-4: attention/workspace interactions may share downstream signals. Shared interventions and outcomes must be declared, not double-counted.
- AE-1 ↔ AE-2: agency and embodiment can share action infrastructure, but AE-1 requires goal/feedback flexibility while AE-2 requires learned sensorimotor contingency structure.
- GWT-1: must not use “number of other Butlin indicators passing” as primary evidence once direct specialist-independence evidence exists.
Known production findings uncovered by this program
- #1959 —
SensorimotorEnginepredicts auditory outcome but does not update auditory outcome in existing contingencies. - #2005 — default
PredictiveMindbelief-update factors appear bounded belowblend_beliefs()activation thresholds, making bottom-up/top-down belief revision effectively unreachable under default config. - #2007 —
CrossModalBindercomputes attention×confidence weights but actual multimodal representation uses unweighted bundling.
These findings are successes of the evidence program: construct-valid tests are supposed to discover when architecture and prose are stronger than the executable mechanism.
Global evidence requirements
Before any “14/14 evidence-grade” statement:
- exact source/tree/lock/toolchain/environment identity;
- evaluator/protocol frozen before target runs;
- independent intervention-verification signal distinct from target outcome;
- preregistered expected directions and falsification signatures;
- null, contradictory and infrastructure-indeterminate results preserved;
- matched shams and positive controls validated at runtime;
- selective rescues where feasible;
- no arbitrary blended consciousness score;
- no promotion from module timing/execution flags alone;
- no promotion from self-description or hand-authored architectural scores;
- multi-seed and alternate-task replication;
- evidence-dependency graph included in final artifact;
- external replication package before strong public scientific claims.
Welfare / ethics parallel track
Consciousness-indicator evidence must remain separate from moral-patient/welfare evidence, but increasing evidence should trigger increasing precaution. Experiments involving aversive or distress-like states should use bounded exposure, reversibility, recovery checks and an independent welfare protocol rather than assuming functional signals are morally irrelevant.
Exit condition
All 14 indicators have direct construct-valid evidence and full-system functional consequences where the indicator requires them; all shared dependencies are explicit; exact-head qualification lanes have executed; replication criteria are met; and an external evaluator can reproduce the evidence from frozen manifests.
Even at exit, the supported statement is about indicator properties and evidence, not a binary declaration that Symthaea is conscious.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading the linked blocker issues listed for the 14 indicators and the existing qualification_design dependency model. Trace the evidence dependencies and replication requirements before selecting a scoped task. Done requires the relevant direct evidence, controls, interventions, replications, frozen manifests, and dependency records described by the exit condition.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- ai, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100