CB-99 — Closing: provenance and "falsify any step"
How to read this page
Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.
The Cookbook is the method carried out step by step: 34 pages of recipes for building a developmental active-inference SIMULATION — a bounded peek at a toy world, never a person. The front matter says that word is never softened under any pressure, so it is not softened here. The recipes run from the molecular and cellular rungs up through metabolism, motor control, perception, language and metacognition, and on to rungs that are still open questions. Around them sit a set of kitchen rules, a shared pantry of engines and primitives, and a second family of recipes about nature itself — rocks, water, air, stars, DNA, ants, whales, bats, humans.
It is for the reader asking what building this would actually take. Each recipe names its ingredients, the order of work, the tests to run at that stage, and the point at which a step stops being something already carried out and becomes something proposed.
Begin with the front matter and then the kitchen rules. Those two pages fix the honest position and the fence labels that every later recipe leans on, and without them the status markers on a recipe are easy to skim past. After that the recipes can be read in any order.
The nature recipes sit slightly apart and should be read that way. They cite outside science — geology, chemistry, biology, astrophysics — and a nature citation is never a UNI gate: those chapters contain zero UNI claims and raise no rung.
What it is not: a claim that the whole ladder has been cooked. The book recommends the complete recipe and, on the same page, labels every rung by its real state — that tension is deliberate and is the thing the book is built around. Where a recipe and the claim ledger disagree, the ledger wins.
Your browser cannot switch reading levels, so the document itself is shown.
Precise — the source document
This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.
What you are building: not another recipe, but the seal on every recipe before it — the provenance trail that lets a stranger trace any figure in this book back to its source row, and the standing invitation to knock the whole thing down.
This is the last page of a literal cook book. Every chapter before it told you which pantry engines to reach for, the exact build steps, the gate to clear, the falsifier that would sink it, and the recorded negatives that sit beside it. This chapter does the one thing a recipe book owes its reader at the end: it shows where the numbers came from, and it hands you the knife.
The constitution you have been cooking under
Six rules bound every page, and they bind this one too. They are not capability claims; none can be raised above its stated role.
- The ledger is the single source of truth. This Cookbook authors against
../encyclopedia/CLAIM-LEDGER.md. Where this recipe and the ledger disagree, the ledger wins and this recipe is wrong. - Calibration moves DOWN only. Authority flows from the measured fact. Wording is calibrated to the measured value, never up — including under urgency. The fence gets louder under pressure, not wider.
- The verdict is the CI bound that excludes the threshold, never the point estimate.
- Negatives are content. The corpus carries 183 published negatives (last ledger snapshot: 882 rows = 350 PASS / 0 FAIL / 183 NEGATIVE / 349 PENDING). The negatives are the credibility, shown front-and-center, never deferred.
- No claim is carded above its source evidence class. Class authority is fixed: A (machine-exact anchor), B (mechanism + operator observation), C (dev-gate / held-out eval), E (test-covered), U (claimed-but-unproven — and "Class U — not claimed" is itself a standing fence).
- The hard fences hold in every chapter. Never AGI / general intelligence / human-level / "understands". Never consciousness / sentience / aware (functional self-awareness may be described only at L8; phenomenal sentience is explicitly disclaimed — not tested, no falsifier offered). Never "active inference demonstrated" (the framing lens only; no AIF loop exists in the Rust crate, and the live UNI.OS loop is a separate reimplementation, not gate-matched). Never "created life" / "digital life" / "measurable awareness" as a claim. Never "beats LLMs" (World C is a count baseline, ~10–15% behind backprop LLMs on char-perplexity by a chosen design trade). Never inflate the Tier-2 synthetic-construction track into capability. No PII; no patent-level math (textbook-level framing only).
Reading the FENCE label
Every recipe closed with one of exactly four words. Carry the vocabulary precisely; it is drawn from the ledger and never inflated.
- proven — a held, sealed, UNI-signed PASS exists (Class A/C), with its falsifier still live.
- designed — a typed spec / signed-in-principle design exists; the build is partial or unverified.
- hypothesized — a stated mechanism or law with no sealed gate yet (Class U, not claimed).
- not-yet-built — no engine, no run, no gate; north-star, hard-fenced.
A designed consult does not become a proven result by being signed. Every SIGNED consult design
folded into this book — L2's build_epistemic_frontier organ and the G6_BUILD_EPI_FRONTIER paired RED
(G6 stays OPEN); L5 Design #3, the Proprioceptive Servo Bridge (not-yet-run); the L9-G1 cavity-correct
gate; the L11-R1 offline-replay gate; the L12 framing bound and checklist; the C10 Dirichlet-first port —
is folded in as designed / not-run. Each raises nothing, and the rung's status is unchanged.
Provenance: where every figure traces
Every quantitative figure, PASS / NEGATIVE verdict, evidence class, and falsifier in this book traces to a row in the ledger:
- the L0–L12 developmental rungs,
- the C0–C14 continuity / embodiment-substrate sub-ladder (engineering evidence, NOT general-AIF evidence — substrate work never implies a science gate is met),
- the M1–M30 method / evidence-constitution patterns,
- and the TA-N12 sharpened north-star bar.
The continuity sub-ladder carries its own first-class recorded negatives — engineering / substrate evidence, never general-AIF evidence:
- C7 — a single-box live OS update is a seconds-long freeze, not zero-downtime; external media legs (RTP/SIP/kernel-mode rtpengine) are NOT preserved across kexec (true zero-freeze needs the second node carrying the platform).
- C8 — the over-compressed 2-modality sensory bottleneck went NEGATIVE on held data; keep all 7 modalities.
- C9 — the EDAIT trade: an exact-discrete active-inference transformer trades fluency for calibration — held-out perplexity ~33 vs a backprop GPT's ~25 (less fluent, but natively online-learning + calibrated). An honest trade, not a win.
- C14 — the multi-tenancy isolation gap: the bare UNI substrate is missing tenant / client / namespace isolation (flat global perms) and temporal decay on learned counts; the product build is the tenant wrapper + decay, not the core.
The ledger itself is the deduplicated merge of 615 extracted claims across twelve source archives,
reconciled through 00-INDEX. Its maintenance rules are the provenance contract for this whole book:
- Authority: every row is carded at the lower of its evidence components and at its strongest source archive. Cross-referenced rungs are merged into one canonical row, not re-listed per archive.
- Append-only: corrections are forward-only (supersede with lineage), never silent edits. Calibration moves wording DOWN only.
- Single source of truth: if prose and the ledger disagree, the ledger wins and the prose is wrong.
- Preprint citation, always fenced: Polzin et al. 2026, Zenodo DOI 10.5281/zenodo.19785799 (MIT) — the mathematical foundation only, an unrefereed working preprint (Layer-1 AI-executable audit complete; Layer-2 human expert review PENDING). Never cited as proof that active inference is the correct theory.
- Science provenance, textbook-level only: Parr, Pezzulo & Friston, Active Inference (MIT Press, 2022). Patent-level UNI math stays private.
A peer-reviewed honesty audit (os-cycles 39–53) already calibrated six headline claims down to the measured value — "cleared on TWO boxes" became ONE box; "the mind survives a patch" became infrastructure continuity only; "5 stacks" became "4 distinct stacks / 5 runs." Those overstatements lived in the headline layer, not in fabrication. This book carries the calibrated figures, never the inflated ones.
The honest position (printed, never softened)
~2 of 11+ developmental rungs earned.
Proven, today, with a sealed, UNI-signed, still-live falsifier: L0 (zygote first-division as exact
discrete Bayes, Class A, float32 tier ~6e-8 — never the f64 tier); L1 (the Cell Lab benchmark,
Class C, with its honest losses on database_flaky, memory_leak, cpu_noisy_neighbor); L3 (the
Karaaslan Heart Lab, Class C, a toy model, not a clinical instrument); L4 (affect-as-precision, Class
C, affect modeled, never felt); L6 (World C, +0.081 nats/char, CI [0.0736, 0.0890], the
one citable empirical PASS — a count baseline, not active inference, not comprehension, not
beats-LLMs); L7 Phase J (+0.105 nats/char, CI [0.0975, 0.1128], always cited beside the
J.attribution_caveat NEGATIVE); L8 (the M1–M11 report card 15/15, functional self-awareness,
phenomenal sentience disclaimed). L5 carries a sealed synthetic PASS (Design #1, +0.092, CI [+0.038,
+0.157]) beside a sealed synthetic NEGATIVE (Design #2, −0.091, CI [−0.134, −0.055]).
World C is the lone winner, and its recorded negatives carry beside it: FIVE within-segment-structure designs went NEGATIVE-with-discriminator (the Phase G char-perplexity Section 0.6(B) bound — no within-segment structure beats tuned MKN-7; char-ppl is a chosen design trade), and two controller families showed the World-C gain is diffuse with nothing to gate (the Phase F active-controller NEGATIVE).
At L7, beyond the paired J.attribution_caveat, comprehension above retrieval is a genuine published
wall — the thrice-NEGATIVE central wall: K≥3 structurally-distinct no-backprop designs fail to beat
retrieval baselines on adversarial comprehension. It is a recorded negative, not a hidden failure.
L2 is a real foraging/crafting driver (+135% / 2.35× tool-crafting, +19% mining) but a recorded negative as a building driver (building went −14%, G4 allostasis never separated); the plateau break (G6) stays OPEN, diagnosed as epistemic starvation, not γ-runaway. L9 and L10 are parked (Class U, designed gates only, nothing claimed). L11 and L12 are not-yet-built (hard-fenced; the awareness question is posed as an open, falsifiable question, never an answer). The L7 char-perplexity / T2 word-grain frontier and the L9–L10 ladder are parked as a ledger-scoped exhausted search envelope: a scoped negative over the tested envelope only — not a universal impossibility result and not an achieved rung. The registered tested K conditions did not reverse the result. Sec-0.6(B) remains unearned / parked.
The whole program is a developmental active-inference SIMULATION: a bounded peek, a toy world, never a person.
Falsify any step
This Cookbook fully recommends the complete build to L12 — assume the full recipe. But every step is labeled by its real status, no step is carded above its source evidence class, no negative is hidden, and no fence is crossed. Roughly 2 of 11+ rungs are earned; the rest are designed, hypothesized, or not-yet-built, and they are marked as such on their own pages.
So here is the knife. Take any proven rung and re-run it on a fresh held-out split with ≥5 disjoint seeds: if the seed-paired bootstrap CI on the margin includes or falls below its threshold, the rung falls. Show World C's baseline was untuned, or collapse Phase J's gain under marker-swap, or pass a self-model task at L8 with a trivial heuristic, and the PASS is gone. Demonstrate mind-tick continuity across a real kernel swap and you discharge the owed Stage-2; show a within-segment structure that beats tuned MKN-7 and the char-perplexity Phase G bound reopens; beat a tuned retrieval baseline on adversarial comprehension with a no-backprop design and the L7 thrice-NEGATIVE central wall falls. Every gate in this book names the exact observation that would sink it. None of them is hidden, and none is exempt.
Falsify any step.
HONEST FENCE — method / proven (constitution) over a ~2-of-11+ position. This closing chapter carries the constitution, the 4-value fence vocabulary, and the honest program position exactly as the ledger states them. It claims nothing beyond what the ledger holds: not AGI, not consciousness or awareness, not "active inference demonstrated", not "created life", not "beats LLMs", not any rung raised by a signed design. UNI remains a developmental active-inference simulation; the awareness question remains open. Where this recipe and the ledger disagree, the ledger wins.
sha256 c1b81782eeb8ee27 — of the original file, so what was ingested stays checkable.
Plain — written for this website, not the source document
The last page of the cookbook does two things a recipe book owes its reader at the end. It shows where every figure in the book came from, and it hands the reader a knife.
The provenance half explains that each number, verdict, evidence class and falsifier — the observation that would sink the claim — traces back to a row in a separate ledger of claims, a list added to and never edited. It was merged from many source archives, under rules that only ever move wording down toward what was measured. It records that an honesty audit already calibrated several headline claims downward, and it names them.
The other half is the invitation. For each step the book calls proven, this page states the exact observation that would sink it. Re-run it on fresh data it has not seen, show a baseline was untuned, or pass a task with a trivial shortcut instead. Nothing is exempt.
Between the two halves it prints the honest position again. This is a developmental simulation — a bounded peek at a toy world, never a person — with about two of eleven-plus rungs earned, and the awareness question left open.
Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is c1b81782eeb8ee27
Clear — written for this website, not the source document
This closing chapter is the seal on everything before it. Rather than adding a recipe, it makes the whole book traceable and then invites anyone to knock it down.
It opens by restating the constitution that bound every earlier page. A separate ledger of claims, added to and never edited, is the single source of truth, and where the book and that ledger disagree the book is wrong. Calibration moves wording only downward toward what was measured, including under urgency. A verdict rests on the bound that excludes the threshold rather than on a best estimate. Negative results are content, shown up front, because they are the credibility. No claim is carried above its source evidence class, and a short set of hard limits holds in every chapter: no general intelligence, no awareness, no demonstrated active inference, no created life, no beating other systems.
The four status words are restated, with one point made very firmly. A design does not become a result by being signed off. Every signed design folded into the book enters as designed and not run, raises nothing, and leaves its rung's status unchanged.
Provenance is the heart of the page. Every quantitative figure, verdict, evidence class and falsifier — what would sink the claim — traces to a ledger row. The rows cover the developmental rungs, a separate continuity sub-ladder that is engineering evidence rather than science, a set of method patterns, and a north-star bar. The continuity negatives are carried first-class. A live update on a single box is a seconds-long freeze rather than no downtime, and some external legs do not survive it. An over-compressed sensory bottleneck lost on data it had not seen. A fluency-for-calibration trade is called honest rather than a win, and the bare substrate is missing isolation and decay. The ledger's own maintenance rules are given too: each row takes the weaker class of its parts, corrections are added rather than edited in silence, and the preprint is always cited as mathematics only and marked unrefereed. It records that an honesty audit already calibrated several headline claims downward, naming each one.
Then the honest position, printed without softening: this is a developmental simulation, a bounded peek at a toy world, never a person, with about two of eleven-plus developmental rungs earned. The chapter lists which rungs hold a sealed pass and prints their limits beside them. A first-division result at the coarser numeric tier. A cellular benchmark carried with its named losses. A heart model that is a toy rather than a clinical instrument. An affect model that is modelled and never felt. A reading result that is a counting baseline and not comprehension. A language result always cited beside its own caveat, and a self-model that is functional only. The motor rung carries a sealed positive beside a sealed negative. Several negatives are set out at length, including a wall where comprehension above retrieval has not been beaten across structurally distinct attempts, and a metabolism result that helps foraging and crafting while making building worse. Two rungs are parked, and two are not built at all.
The last section is the knife, and it is specific. Re-run a proven rung on fresh unseen data with several disjoint seeds and watch whether the interval still clears its threshold. Show a baseline was untuned. Collapse a gain by swapping markers. Pass a self-model task with a trivial rule. Show continuity across a real kernel swap and an owed stage is discharged. Every gate names the observation that would sink it, and none is hidden or exempt.
Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is c1b81782eeb8ee27