Wiki · Evidence & Verdicts
Forage RED — pre-registration (Cure-1, isolated drive test)
How to read this page
Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.
Eighty-seven dated pages: receipts, pre-registrations, handoffs, validation records and review verdicts. A receipt is written at the moment a piece of work was checked. It names what was claimed, the commit and the seed, what was actually run, and the outcome in one of a small set of controlled words. Then it names what the work did not achieve. That last part is what makes it a receipt rather than an announcement. A pre-registration is the same discipline run in advance: the conditions that would count as a pass and the conditions that would falsify the claim are written down before the run, so neither can be adjusted once the numbers arrive.
That is why so many small dated stubs are an audit trail rather than noise. No one of them is meant to be a good read. The value is in the sequence and in the dates, because you can watch a prediction be registered, then the run happen, then the verdict land — sometimes against the prediction. Pages here record a falsified result, a rejected fix, a retracted overclaim, and a green receipt that turned out not to be reproducible from the commit that carried it. A record that carried only successes would be worth a good deal less than this one.
A gentle way in is to read a pre-registration first, so the shape becomes familiar, then a result page, then one of the corrections. This section sits off the main navigation on purpose: it is the record you check the rest of the site against, not the place to begin.
What it is not: documentation, and not a summary. Nothing here has been tidied in hindsight. Every entry reads as of its date, a later entry may overturn an earlier one, and the presence of a page is not a claim that its result stood.
Your browser cannot switch reading levels, so the document itself is shown.
Precise — the source document
This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.
Pre-registered BEFORE the run (Lab Protocol §pre-registered RED gates). Owner go-ahead recorded 2026-07-11.
Runner: runs/forage_red.exs. Branch lab/ozone-life-uni-hard-science @ c016171.
Question (first rule, C11)
Does the epistemic (novelty) drive CONSTRUCT emergent hunting that a driveless twin lacks? This is the isolated Cure-1 test that must have a recorded verdict BEFORE the nursery/training bundle runs on top.
Design
Two arms, both FRESH (untrained), same prey-stocked world, same MILD developmental runway (metab_scale
SCALE=0.5) so neither starves before the drive can act — the only difference is novelty_gain:
- ON = kin 72,
Genome.nursery(0.3, 0.5)— forage novelty ON + equal runway. - OFF = kin 73,
Genome.nursery(0.0, 0.5)— forage novelty OFF + equal runway (the control twin).N_PER_ARM=3,SOAK_SEC=2700(45 min),WARMUP_SEC=180. Isolation: distinct kin (72/73) + distinct memory dir, on the shared idlemc-server(streamed colony DOWN, 0 players) — the established RED convention (motor/curiosity/metabolism). Peaceful + day-locked (isolates foraging from combat/night death). Prey are summoned live animals the bot must hunt; ZERO calorie gives (the launcher RAISES on give/item/clear/xp).
Mechanism under test (reward-free)
interoceptive-depleted → L2 :forage → prey-orient C → :attack under-sampled ⇒ transition-novelty W_b
(ON only) makes it worth TRYING → world-earned kill (body.js collectDrops) → Dirichlet B learns
attack→has_food. OFF has no exploration pressure to try the strike.
Gates (pre-registered — no post-hoc retuning)
PASS — the ON arm MATERIALLY out-forages the OFF arm on ≥3 of these 4, INCLUDING the mechanism (atk_food):
atk_food(learnedpb[:attack]→has_foodmass) ON > OFF (the mechanism — REQUIRED for PASS).food_seen(probes with world-earned meat in hand) ON > OFF.refills(energy up-jumps ≥0.15 = a hunt→eat cycle) ON > OFF.alive_frac(survival fraction) ON ≥ OFF.
FALSIFIES — ON does NOT out-forage OFF (drive inert): atk_food ON ≤ OFF, or no material difference on 2+.
VOID (re-run) — both arms die out before warmup (window too harsh, alive_frac ≈ 0 both) → raise SCALE
or lengthen the soak and re-run; OR any gives>0/summons>0 leaked (structural — the launcher guards it).
Claim fence
A PASS demonstrates the reward-free BEHAVIOUR "the novelty drive constructs the hunt" — necessary-not-sufficient,
zero evidential weight for awareness / hunger-as-experience / life. Every store/count/pb mass is a MODEL
VARIABLE. Survival = in-world persistence only.
After this verdict
Recorded PASS → proceed to the nursery→pure-world QA gate (runs/nursery_forage_gate.sh) under owner go-ahead;
ping the Producer for G2 only after a TRAINED brain forages+survives a scaffold-free pure world.
FALSIFIES → the drive does not construct foraging; do NOT proceed; diagnose (fix forward, no rollback).
RESULT — Run 1 (recorded 2026-07-12, live on mc-server, container uni-colony-forage)
VERDICT: WITHHELD (inconclusive on the drive — the emergent hunt loop did not engage for EITHER arm).
Evidence: ~uni/.claude-evidence/forage_red/run1_20260712T000630.log on the lab box. Run was clean —
gives=0 summons=0 colony_ok=true c_ok=true both arms (no manna leaked; the isolation held).
| arm | survived | mean_energy | died by | attack | eat | food_seen | refills | atk_food |
|---|---|---|---|---|---|---|---|---|
| ON (kin72, novelty 0.3) | false | 0.13 | ~t24 (~6 min) | 0* | 0 | 0 | 0 | 0.25 (untrained prior) |
| OFF (kin73, novelty 0.0) | false | 0.17 | ~t24 (~6 min) | 0* | 0 | 0 | 0 | 0.25 (untrained prior) |
Death curve: live 6/6 (t0,t8) → 5/6 (t16) → 0/6 (t24). All 6 bots starved in ~5–6 minutes. Body logs show
only CRAFT/spawn categories — zero attack/hunt/collect/eat activity from either arm.
(*attack/eat/atk_food are read from the dead final sample — an aggregation bug — but food_seen/refills,
counted over ALL live samples, are a true 0, and the body logs confirm no hunting occurred while alive.)
Why WITHHELD, not FALSIFIES-of-the-drive: the RED could not fairly test the drive because NEITHER arm foraged — you cannot distinguish "the novelty drive is inert" from "the hunt loop never engaged for anyone." Two coupled causes:
- Window too harsh (the pre-registered VOID condition, ~met):
metab_scale 0.5→ death in ~5 min, far too fast for learned foraging (many hunt trials) to emerge. A gentler runway is required. - The forage policy did not engage hunting at all (the deeper, real finding): while alive AND hungry, the
bots gathered/crafted (wood/tools — all
CRAFT no-recipe, they had no wood) instead of orienting to and striking prey. The interoceptive-hunger →:forage→ prey-hunt path is being out-competed by the default gather/craft behaviour; the novelty drive did not redirect it to the strike.
Fix-forward (no rollback — follow EFE/VFE):
- Instrument a diagnostic re-run: log per-bot chosen action + L2 situation/context + whether
:attackis ever selected + prey range — to confirm WHY no hunting (leading hypothesis::forage's prey-orient C is outweighed by the phase-0 wood/build preferences when depleted). - Fix
forage_red.exs: aggregate attack/eat/atk_food as max-over-LIVE-samples (not the dead final sample); stop stocking + exit early once all bots die (this run wastefully summoned ~180 mobs post-death). - Likely FE gap to address (gated + reviewed): when
energy_reserveis critical, the interoceptive hyper-prior must dominate the policy — suppress the gather/craft pull and amplify prey-orient + attack so hunger actually drives hunting. This is the honest next cure; it needs its own design/review before deployment. - Only after the hunt loop is shown to ENGAGE (bots attack prey + secure meat) does a gentler-window survival comparison (novelty ON vs OFF) become meaningful.
Claim status: the emergent-forage mechanism did NOT produce live hunting/survival in this configuration. This is a receipt of a real negative result — the live embodiment falsified the assumption that gaps 1+2 + novelty would yield hunting out of the box. Do NOT proceed to the nursery/QA gate or ping the Producer for G2.
DIAGNOSIS (2026-07-12) — two stacked failures, isolated
Offline decision probe (runs/probe_forage_decision.exs, MC.step/2 sampled per scenario): the L2 wiring is
CORRECT (hungry ⇒ situation=2 depleted ⇒ :forage), but the policy overwhelmingly picks :eat (7–12 of ~13)
and :attack ~1 in every hungry scenario — even with prey ahead and an EMPTY inventory. Root: designer.ex
transition(:emptying,…) gives the :eat column a FILL matrix (energy +2 bins) and every other action a DRAIN
matrix, so :eat is the UNIQUE energy-raising action in the model, seeded hard (pb_seed 50), unconditioned on
inventory food. A depth-5 planner satisfying the reserve-C therefore always prefers :eat; :attack only ever
drains ⇒ never pragmatically chosen. The appetitive/consummatory gap: eating is EFE-minimal but useless; food
acquisition (hunt) is never valued. (POLICY layer.)
Live instrumented diag (runs/forage_diag.exs, per-bot action/context sampling): max_inv_food=0 for ALL
bots — even one that struck 11× got zero food. Root: body.js doAttack only swung at an entity ALREADY within
reach+crosshair and never CLOSED the distance (unlike mineTree, which walks up to the log). Side/far prey
(prey=2/3) was never struck. (MOTOR layer — the BINDING constraint: no food is physically possible.)
FIX 1 — hunt motor (ff57a5a, body-only, CONFIRMED WORKING)
doAttack is now a hierarchical hunt motor symmetric with the validated mineTree: resolve nearest prey/hostile
(ACT-path), close the gap (≤6 forward+hop steps re-facing), strike ≤4× until it drops, collect the meat. No
generative-model / invariant / gating touch; the brain still CHOOSES :attack and LEARNS prey+:attack→has_food.
DIAG-2 verdict (motor isolation, scale 0.2): CONFIRMED — emergent survival-by-hunting is now POSSIBLE.
UNI-72-1 (novelty ON) hunted, collected max_inv_food=12, spent most of the run CALM (sit=0) at FULL
energy (e=5 T180→T720) — hunt→kill→collect→eat→stay-fed, live, world-earned, zero gives. Evidence:
~uni/.claude-evidence/forage_red/diag2_bodyfix_*.log. The other 3 bots still starved (inv_food=0: walked-
forever / spun / struck-without-killing) ⇒ hunting is now POSSIBLE but UNRELIABLE — the remaining POLICY gap.
FIX 2 — policy cure (in adversarial design, workflow wf_0cf41964-463)
Make hunger reliably drive CHOOSING to hunt over spamming :eat-on-empty (the FEP-faithful appetitive/
consummatory split), gated + default/homeostat_colony byte-identical, lab-team-reviewed before any FE code.
Deploy + RED AFTER the review, one cure at a time.
sha256 ae2c15c9191710b5 — of the original file, so what was ingested stays checkable.
Plain — written for this website, not the source document
A pre-registration — the conditions written down before the run — with its result, its diagnosis and two fixes attached underneath, all on one page. The question was whether one drive constructs hunting that a twin without it lacks. The verdict is withheld rather than negative, because neither arm hunted at all, so the comparison could not be made. Two causes are separated: the window was too harsh, and, more tellingly, the bodies gathered and crafted instead of hunting while hungry. A later probe found the deeper reason, and the fix proceeds one cure at a time.
Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is ae2c15c9191710b5
Clear — written for this website, not the source document
One page carrying four things in order: a pre-registration written before the run, the result, a diagnosis, and the fixes that followed. The order matters, because the verdict is not what the pre-registration expected.
The design is two arms, both starting untrained in the same prey-stocked world with the same gentle runway, differing in one setting only. The isolation is described in detail. There are distinct lineages and separate memory, and a peaceful day-locked world so feeding is separated from fighting. There is prey that must actually be hunted, and a launcher that raises an error if anything is given rather than earned.
The gates are pre-registered with no post-hoc retuning. A pass needs the treated arm to out-forage the control on most of four measures and specifically on the mechanism measure, which is required. What would show it wrong is the drive being inert. And a void condition is written in advance for the case where both arms die before the window even opens.
The result is withheld. Every body starved within minutes, and the logs show no hunting at all from either arm. The page explains why that is withheld rather than a refutation of the drive: you cannot separate an inert drive from a loop that never engaged for anyone. Two coupled causes are then distinguished. The window was too harsh, close enough to the pre-registered void condition to be named as such. And the deeper finding is that while alive and hungry the bodies gathered and crafted instead of orienting to prey. An aggregation bug in the harness is flagged, with the reason the conclusion survives it.
The fix-forward list is explicit that nothing is rolled back. Instrument a diagnostic re-run to find out why, and fix the harness aggregation so it stops working after everything has died. Only then consider a change to the policy, which would need its own design and review.
The diagnosis section isolates two stacked failures. Offline, the higher-level wiring is correct, but the policy overwhelmingly picks eating over attacking, because in the model eating is the only action that raises energy and it is not conditioned on actually holding any food. A planner satisfying the preference therefore always prefers it, while attacking only ever drains. The page names this the appetitive and consummatory gap: eating is optimal and useless, and acquiring food is never valued. Live, no body ever obtained food, because the motor only swung at something already within reach and never closed the distance.
The first fix is motor-only and reported as working. The attack becomes a hierarchical hunt that resolves prey, closes the gap, strikes and collects, while the brain still chooses and still learns. One body then hunted, fed itself and held full energy for most of the run, world-earned with nothing given. The others still starved, so the honest summary is that hunting is now possible but unreliable, which leaves the policy gap standing. The second fix is in adversarial design and must be reviewed before any code touching the model is written.
Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is ae2c15c9191710b5