Wiki · Evidence & Verdicts
Phase-2 metabolism RED — committed receipt + honest verdict (2026-07-11)
[redacted: category] — 1 private address. Nothing else was altered. The document is otherwise exactly as it is written in the repository, and the sha256 below is of the original, so what was ingested stays checkable.How to read this page
Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.
Eighty-seven dated pages: receipts, pre-registrations, handoffs, validation records and review verdicts. A receipt is written at the moment a piece of work was checked. It names what was claimed, the commit and the seed, what was actually run, and the outcome in one of a small set of controlled words. Then it names what the work did not achieve. That last part is what makes it a receipt rather than an announcement. A pre-registration is the same discipline run in advance: the conditions that would count as a pass and the conditions that would falsify the claim are written down before the run, so neither can be adjusted once the numbers arrive.
That is why so many small dated stubs are an audit trail rather than noise. No one of them is meant to be a good read. The value is in the sequence and in the dates, because you can watch a prediction be registered, then the run happen, then the verdict land — sometimes against the prediction. Pages here record a falsified result, a rejected fix, a retracted overclaim, and a green receipt that turned out not to be reproducible from the commit that carried it. A record that carried only successes would be worth a good deal less than this one.
A gentle way in is to read a pre-registration first, so the shape becomes familiar, then a result page, then one of the corrections. This section sits off the main navigation on purpose: it is the record you check the rest of the site against, not the place to begin.
What it is not: documentation, and not a summary. Nothing here has been tidied in hindsight. Every entry reads as of its date, a later entry may overturn an earlier one, and the presence of a page is not a claim that its result stood.
Your browser cannot switch reading levels, so the document itself is shown.
Precise — the source document
This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.
This is the committed receipt that docs/specs/metabolism.md §16 and docs/DEEPENING_PLAN.md were holding
open as PROVISIONAL. The verdict below is the corrected reading after adversarial verification (4
independent skeptics, workflow wf_61b43ee9-68d); it is materially more conservative than the pre-verification
draft. Where this receipt and any earlier prose disagree, this receipt wins.
Provenance (raw evidence)
- Source: the live collector output on the rootless colony box (
uni@[redacted: private-address]),~/.claude-evidence/phase2/phase2_red.jsonl, pulled 2026-07-11. - Committed here gzipped:
phase2_metabolism_red.jsonl.gz(raw 18,293,470 B, 26,244 rows, schemaphase2_red/1) +phase2_metabolism_red.heartbeat.jsonl.gz(2,187 heartbeats). - SHA256 (raw):
957aafd6f406f7852376a6a49a9c18d12102eb16cee87a1a55f74c1ec86a2b27(phase2_red.jsonl);212c6d31b9d31351f1d0693c724b9df59e6fdfb3e0dd6dd0324cfe8daf8391b3(heartbeat.jsonl). - SHA256 (gz committed):
42bb66129ab972d8e0e31a709414d428ec28f047bffaa89f542639ec82c90a05(phase2_red.jsonl.gz);fe5290e7ced43f06113482935c318a74e7277cdd69da80cffd7e1666bd12b11c(heartbeat.jsonl.gz). - Instruments: launcher
runs/metabolism_lineage.exs; collectorops/phase2_collector/{collect.exs, collector.container,uni-collector.{service,timer},harness_task.md}(already tracked).
Design (what ran)
Paired RED, N=6/arm, ~15 days (2026-06-25T22:31Z → 2026-07-11T02:30Z), container
uni-colony-metabolism (started 2026-06-26 01:23Z) against mc-server (Paper, seed 8675309):
- kin-12 = TREATMENT =
Genome.metabolism_primary()(the:metabolismorgan ON), pinnedphase: 1. - kin-13 = CONTROL =
Genome.default()(organ OFF), pinnedphase: 1. Same code/world/body; differs ONLY in the organ. novelty_gain = 0.0in BOTH arms (the epistemic term was OFF everywhere by construction).- Heartbeat:
node_attached+rcon_connectedtrue throughout;rows_written=12/cycle; 100% probe/rcon health.
The numbers (final per-UNI, RCON-authoritative)
G6 plateau-break metric = placed_used_total + distinct_mined_types (pre-registered, above control).
| arm | per-UNI placed_used_total |
Σ placed | Σ distinct_mined | Σ stone+cobble mined | mean action_entropy |
|---|---|---|---|---|---|
| TREATMENT (kin-12) | 8, 10, 20, 12, 14, 8 | 72 | 13 | 83 (UNI-12-3 alone; others 0) | 1.089 |
| CONTROL (kin-13) | 16, 2, 14, 3, 34, 14 | 83 | 15 | 64 (13-5=53, 13-6=10, 13-1=1) | 1.152 |
Cumulative placed_used_total by day (freeze): treatment 7→70 (06-26)→72 by 06-27, flat for 14 days;
control 9→76→…→83 by 07-01, flat for 10 days. Nobody in either arm mined cobblestone or built shelter
(0/12); distinct_block stayed 1–3.
Verdict — split (corrected, adversarially verified)
1. G6 plateau-break (the pre-registered PASS gate: treatment must EXCEED control; behavioural target =
reach cobblestone/shelter) → FAIL. This is the only firm, noise-immune conclusion. Treatment did not
exceed control on placed_used (72<83) or distinct_mined (13<15), and 0/12 UNIs in either arm reached
cobblestone or built shelter. The cure did not break the plateau. (A narrow FAIL of the directional gate +
the absolute-target miss — both robust to the noise below.)
2. The metabolism HYPOTHESIS itself (did the organ help/hurt?) → WITHHELD. Two independent reasons:
- Arms are statistically indistinguishable at N=6. Control
placed_usedranges 2–34 (SD≈11.6, ~2.5× the treatment SD); Welch t≈−0.36 (p≈0.73); the difference 95% CI ≈ [−14.3, +10.6] straddles zero and does not exclude either threshold. Per the protocol's own rule ("verdict = the CI bound that excludes the threshold, never the point estimate"), the 72-vs-83, 13-vs-15, and entropy 1.089-vs-1.152 gaps are within noise — "treatment did worse / explored less / froze harder" is not supported. The sign is not even robust: leave-one-out on control UNI-13-5 (=34) flips it to treatment-leads. - Organ activation is UNVERIFIED. The pre-registered mechanism falsifier — G5b, the action-severed
energy-axis twin — was never passed (
metabolism.md§16 says so). The RED collected only G6 behavioural RCON metrics; there is no energy-posterior / twin-survival receipt showing the internal energy/satiety store depleted, refilled, and modulated action selection in the deployed containers. We therefore cannot say metabolism "failed" — only that this run licenses no metabolism claim.
Struck from the earlier draft as over-reach (do not restate):
- ❌ "metabolism produced a sustained foraging/mining homeostat that is itself a new plateau" — the organ-free control froze identically, so the freeze is baseline / a shared world-or-observation-bin ceiling, not an organ effect. A flat line is not a regulated set-point.
- ❌ "consistent with epistemic starvation (no epistemic drive)" as mechanism —
novelty_gain=0in both arms is a fixed background condition of the whole rig, not a finding of this RED; this run cannot adjudicate the epistemic drive. - ❌ the stone flip (treatment 83 > control 64) as any kind of win — reported here for completeness, then rejected: it is entirely one UNI (12-3=83, the other five treatment UNIs = 0), not the pre-registered metric, and never converted to placement or the shelter target. A per-arm sum dominated by n=1 is a lucky forager, not a plateau break.
Claim fence (binding)
Every number here is a behavioural/model observable with zero evidential weight for awareness / experience / life. A null / indistinguishable behavioural result carries even less such weight than a positive one: this run demonstrates no distinctive behaviour attributable to the organ, let alone experience.
What this run does — and does NOT — license for the next design
- It licenses: "the metabolism organ, as deployed at curriculum-phase-1 with novelty off, did not demonstrate a plateau-break, and both arms hit a shared low ceiling."
- It does NOT license: "metabolism failed," "metabolism causes a homeostat," or "the plateau is epistemic starvation." The Track-B reframe (true-signal vision + a natural epistemic drive) remains a hypothesis this RED neither confirms nor refutes. The shared arm-independent freeze is itself a signal that a world/observation-bin ceiling (task affordance) may be as load-bearing as any drive deficit — a world-ceiling control ("can ANY configured agent reach cobblestone in this world?") is now a prerequisite.
Blockers to upgrade this from WITHHELD to a real metabolism verdict
- ✅ Commit this receipt from the lab-box JSONL (done).
- The G5b energy-axis twin / energy-posterior receipt proving the organ modulated action in the live run.
- Per-UNI CIs (not arm means), given single-UNI dominance.
- A world-ceiling control separating task-affordance failure from any drive failure.
Adversarial verification
Verdict corrected via workflow wf_61b43ee9-68d (4 lenses: metric-validity, statistical-robustness,
confound-hunt, claim-fence). The pre-verification draft ("PARTIAL/FAIL; metabolism homeostat; epistemic
starvation") was rejected as (a) adjudicating on point estimates the protocol forbids, (b) imposing a
mechanism narrative on a null with an unverified organ, (c) using an invalid verdict token. This receipt is
the surviving verdict.
sha256 32c763baee508567 — of the original file, so what was ingested stays checkable.
Plain — written for this website, not the source document
A committed receipt — the file recording what was run — with a split verdict, and the split is the honest part. The behavioural gate, written down before the run, fails: the treated group did not exceed the control, and nobody in either group reached the target at all. But the hypothesis itself is withheld rather than refuted, for two reasons given plainly. The two groups are statistically indistinguishable at this size, and the organ under test was never shown to be doing anything. The page also strikes three claims from its own earlier draft as over-reach, and says not to restate them.
Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 32c763baee508567
Clear — written for this website, not the source document
A receipt, the file recording what was run, that replaces a provisional reading and says so: where it and any earlier prose disagree, this one wins. The verdict was corrected after adversarial verification and is materially more conservative than the draft it replaces.
The provenance section is unusually complete. Where the raw data came from, how many rows it holds, the digests of the raw and the committed forms, and which instruments produced it. The design is a paired comparison over about two weeks, with a small number per arm, in the same world with the same code. The arms differ only in whether one organ is switched on, and a third term is switched off in both by construction.
The numbers are given per body rather than only as arm totals, which turns out to matter. The measure, fixed before the run, is stated, both arms froze early and stayed flat for many days, and nobody in either arm reached the behavioural target.
The verdict then splits. The behavioural gate fails, and that is called the only firm, noise-immune conclusion, because the treated arm did not exceed the control on either component and the absolute target was missed by everyone. The hypothesis about the organ is withheld for two separate reasons. The arms are statistically indistinguishable at this size, with an interval that straddles zero, and the protocol's own rule forbids adjudicating on point estimates; the sign is not even stable, since dropping one body from the control flips it. And activation was never verified, because the mechanism check written down in advance had not passed and this run collected only behavioural measures.
Three claims are then struck from the earlier draft, marked as not to be restated. That the organ produced a sustained homeostat, since the arm without the organ froze identically, so the freeze is a shared ceiling rather than an effect. That the result fits a particular mechanism, since the relevant setting was a fixed background of the whole rig rather than a finding of this run. And a favourable-looking number that turns out to be one body alone, on a measure that was not the one written down in advance.
The claim fence, the stated limit on what may be said, adds a point worth keeping: a null result carries even less weight for questions about experience than a positive one would.
The page closes by separating what the run licenses from what it does not, and lists what would be needed to upgrade the withheld verdict. That list includes a mechanism receipt, per-body intervals rather than arm means, and a control asking whether any configured agent could reach the target in this world at all.
Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 32c763baee508567