N2 - Same math, many scales
How to read this page
Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.
The Encyclopedia is the UNI method written out as a reference work: 39 pages, arranged in wings, setting out what the programme is attempting and why it is built the way it is. This is where the ideas are explained in order and in prose, rather than as code, as runbooks, or as dated receipts.
Every chapter is authored against two ledgers and never ahead of them. One records what UNI has built, and the evidence class of each claim. The other records nature's own regularities, kept separate on purpose. That way a fact about biology is never quietly reused as a fact about the software. Where a chapter and a ledger disagree, the chapter is the thing that is wrong. Every chapter closes with an invitation to falsify it, and a recorded negative is published beside the result it qualifies rather than after it.
Read "How to read this work" first. It is the evidence constitution: the classes, the four ledger states, and the rule that a finished chapter is not the same as a working system. Then the calibration ledger, which carries the figures every other chapter is required to use.
What it is not: a description of a person or of a mind. The programme calls itself a developmental active-inference simulation, a bounded peek into a toy world, and its own index prints how much of the developmental ladder has actually been earned — roughly two rungs out of eleven or more. It is also not a report of what is running today. For what ran, and when, go to the evidence record.
Your browser cannot switch reading levels, so the document itself is shown.
Precise — the source document
This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.
The honest position of this chapter is narrow and load-bearing: the UNI program runs one active-inference engine, and the same engine is shown acting across very different domains, from a single cell's viability through cardio-renal homeostasis to cognitive precision-weighting. That is an organizing claim about the engine, not a claim that any single scale has been solved beyond its own registered gate. "Same math, many scales" is a statement about reuse and parity, audited mechanically; it is never a statement that the program has earned more than it has. The whole program remains a developmental active-inference SIMULATION (a bounded peek, a toy world, never a person, never a mind), with the honest standing position printed in full: ~2 of 11+ developmental rungs earned.
One engine, shown across scales
The unifying thread is architectural, and it is verifiable. Each public lab is a single self-contained HTML page with an inline JavaScript engine, mirrored by the same physics expressed once in canonical TypeScript under api/_lib/worlds/ (for example service_cell.ts, cardio_renal.ts, trauma_loop.ts), and the two are pinned together by a *_parity.ts test that fails the build if the page engine and the canonical engine ever diverge. New labs are built by duplication: copy a working lab, swap only the observation or generative model, keep the variable names identical so the downstream engine code (buildAMatrix, getObsIdx, the expected-free-energy planner) is reused verbatim, and verify the diff is "exactly one nav line per untouched page." This is the inline-engine + canonical-TS + parity-test triad (ledger method row M28, class method), proven twice by construction: the Echo lab is the Precision lab with only the observation model swapped (a range-2 echolocation ray-cast over a bit-identical 64-observation space), which demonstrates the chapter's recurring teaching point that observation model is not the engine.
This triad is a discipline, not a capability. M28 carries no empirical falsifier of the form "the math is right"; its only falsifier is a governance failure: a parity test that does not fail the build when the inline and canonical engines diverge, or a lab duplication that secretly needed engine-code changes. A companion honesty discipline, the framing_guard / cell_framing_guard.ts test (method row M29), fails the build if framing copy, the DOI, accessibility, or the "does not reproduce Rao's method" labeling regresses, keeping the marketing claims calibrated mechanically rather than by good intentions. Cited at their real weight, M28 and M29 buy exactly two things: the demonstrations cannot lie about their math relative to the single canonical engine, and they cannot quietly inflate their own framing. They buy nothing about whether any scale is solved.
The scales, each carded at its own gate (with its travelling negative)
Cellular viability (L1.1, class C). The Cell Lab is an open, pre-registered falsification benchmark: a hidden 216-state service cell, observation-only controllers, a RecoveryScore, a bootstrap 95% CI on the median paired difference that must exclude 0, and eight honesty fences plus a framing_guard test. UNI tops the leaderboard on most modes. That PASS travels, always and inseparably, with its recorded NEGATIVE (L1.2, class C): UNI honestly loses on three modes, shown at the top of the live leaderboard, not hidden: database_flaky (a rule-based SRE controller wins, 0.803 vs 0.759), memory_leak (a neural net wins, 0.810 vs 0.740), and cpu_noisy_neighbor (a neural net wins, 0.824 vs 0.749, where even the UNI-versus-random margin is not significant). The calibrated phrase is "good but not sovereign." The losses are content; citing the leaderboard win without them is an overclaim.
Physiological homeostasis (L3.1, class C, with the Heart-Lab engine ticket OAS-710-T3 at class E, 15/15). The Karaaslan cardio-renal Heart Lab re-expresses a reduced clinical model (the RSNA -> MAP -> sodium/volume loop) in active-inference language, so that a heart attack is framed as a failure of the prediction loop (maladaptive priors, miscalibrated precision, a damaged generative process). The same engine, a different generative model. Its falsifier is operable: the Heart-Lab predictions diverge from the Karaaslan reference beyond the pre-registered tolerance, or fail parity against the canonical TS engine. The travelling fence is mandatory and is built into the artifact: this is not a clinical tool and not a diagnostic instrument. The resemblance to clinical reality is, in the lab's own copy, "an interpretive act, not a measurement," tagged [HYPOTHESIS]; the heart-attack-as-prediction-loop-failure framing applies to the toy model, not to clinical reality.
Cognitive precision-weighting (L6.3, class E). The POMDP-maze Precision lab, the echolocation Echo lab, and the public Precision Lab expose one engine through three precision knobs: gamma_a (sensory/likelihood), gamma_b (transition), and softmax_temperature (policy), sweeping a 2D bifurcation map into distinct behavioral regimes. The math is ported verbatim from the verified engine, with exactly one disclosed extension (a goal drive). The class here is E (test-covered, parity-tested) and never higher: the falsifier is that the tsx/pytest parity tests between the inline-JS engine and the canonical TS/Python engine diverge, or the bifurcation map fails to reproduce the regime boundaries. A test-covered lab is a faithful demonstration of the engine, not a held-out capability PASS.
The cognitive scale carries the program's flagship empirical bound, and it must travel here too. The one citable empirical capability PASS at this scale, World C (a no-backprop COUNT reader beating a tuned MKN-7 baseline by +0.081 nats/char), is explicitly a count-baseline win, not active inference, not comprehension, not "talking," and not a result that beats LLMs (the program runs roughly 10 to 15 percent behind backprop LLMs on char-perplexity by a chosen design trade). The same scale records its paired NEGATIVE bounds: Phase G (L6.4), where five structurally-distinct within-segment designs were all NEGATIVE-with-discriminator, and Phase F (L6.5), where the World-C gain proved diffuse with nothing to gate. "Same math, many scales" never licenses reading any one scale's demonstration as a solved general capability.
The design-only fence
One referenced source must be fenced hard. The activeinference BEAM workbench (Elixir/Jido/Phoenix, discrete-time only) is the project's most architecturally substantive design seed, and its patterns (runtime-derived UI, deterministic replay, single-source-of-formulas provenance, pure-Jido agents) are recorded as method row M27, class method. But the curated record is explicit: this archive carries zero recorded execution evidence (no transcripts, no build, no test run, no replay, no ledger PASS/FAIL/NEGATIVE/PENDING). M27 is "Recorded DESIGN patterns, the underlying build is unverified." Its patterns must never be presented as results. They are a charter, not a capability.
What is NOT claimed in N2
- Ceiling: It is NOT shown that the engine being one engine means any scale is solved, nor that running across scales is itself evidence of general intelligence, comprehension, awareness, or a demonstrated active-inference loop. The most we claim is the exact, in-class statement: one canonical engine is reused across the cellular, physiological, and cognitive labs, with parity enforced by build-failing tests (M28,
method); cellular viability tops most modes of a pre-registered benchmark and loses three (L1.1 C / L1.2 C); the cardio-renal and precision labs are faithful, test-covered demonstrations of that engine (L3.1 C / engine ticket E; L6.3 E), not clinical or capability claims. - Fences engaged: Red line 3 ("active inference" is the framing lens only; no AIF loop is demonstrated by a parity-tested lab); red line 5 (never "beats LLMs"; World C is a COUNT baseline, ~10-15% behind backprop LLMs); red line 7 (never raise a claim above its source class: M28/M27 stay
method, L6.3 stays E, never recarded as a held-out PASS); red line 8 (engineering and parity discipline does not imply a science gate is met); red line 10 (no PII, textbook-level framing only); red line 12 / vocabulary discipline (textbook-level naming, cite Parr/Pezzulo/Friston). The L12 north-star fence stands untouched: nothing here is "created life," "measurable awareness," or "conscious." - Negatives that travel with this claim (cite alongside, never strip): L1.2 (the three Cell-Lab losses:
database_flaky0.803 vs 0.759,memory_leak0.810 vs 0.740,cpu_noisy_neighbor0.824 vs 0.749, last not significant vs random); the L3.1 "not a clinical tool, not a diagnostic instrument"[HYPOTHESIS]fence; the L6.4 Phase-G five-design negative bound and the L6.5 Phase-F diffuse-gain negative that bound the cognitive scale; and the M27 zero-execution-evidence flag on the BEAM workbench. - Parked / owed: No sign-to-park is owed by this chapter itself (M28, M27, L1.1, L3.1, L6.3 are landed or fixed-class rows, not parked frontiers). The wider program owes the parked items those scales feed (the L7/T2 char-perplexity envelope, the L9-L12 rungs); N2 does not discharge any of them and does not raise them. The
activeinferenceworkbench owes per-epic execution evidence (E0-E7) before any of its design patterns may ever be cited as a result. - One-line honest summary a skeptic could not dispute: UNI reuses a single parity-pinned active-inference engine across the cellular, physiological, and cognitive labs, and that organizing fact (which already loses on three benchmark modes and is bounded by recorded negatives at the cognitive scale) is the entire claim: no scale is solved beyond its own gate, and the one design-only workbench cited here has no execution evidence at all.
Falsify this
The lead falsifier is operable and mechanical: if a parity test (*_parity.ts, run via npx tsx, or the tsx/pytest cognitive-lab parity tests) does not fail the build when the inline page engine and the canonical TypeScript/Python engine diverge, or if any lab duplication is shown to have required engine-code changes rather than only an observation/generative-model swap, then the "same math, many scales" organizing claim is false: the labs would not, in fact, be running one shared, drift-protected engine. Secondary falsifiers per scale: the Cell-Lab recorded losses fail to replicate; the Heart-Lab predictions diverge from the Karaaslan reference beyond pre-registered tolerance; or the precision bifurcation map fails to reproduce its regime boundaries.
Sources
- Ledger rows: L1.1, L1.2 (Cell Lab + honest losses), L3.1 (Karaaslan Heart Lab + OAS-710-T3 engine ticket, class E 15/15), L6.3 (Precision/Echo/Public precision labs, three knobs), L6.4 / L6.5 (Phase-G / Phase-F cognitive-scale negative bounds), L6.1 (World C, the count-baseline ceiling, CI [0.0736, 0.0890]), M28 (inline-engine + canonical-TS + parity-test triad), M29 (framing-guard honesty-as-a-test), M27 (BEAM workbench design patterns, build unverified) -
CLAIM-LEDGER.md. - Front matter: FM-1 (Evidence Constitution), FM-2 (A-U rubric), FM-3 (red lines + forbidden-phrasings, UNI-GPT consult 2026-06-27 SIGNED), FM-4 ("what is NOT claimed" template) -
MASTER-PLAN.md, spec section N2. - Narrative grounding (PII-redacted):
curated/uni-precision-digest.md(the lab roster, the triad, the Cell-Lab benchmark and recorded losses, the Karaaslan basis),curated/worldmodels-digest.md(the verified engine, the three precision knobs, the bifurcation map, the[HYPOTHESIS]honesty fences, the textbook math provenance),curated/activeinference-digest.md(the BEAM workbench, design-only, zero recorded execution evidence). - Math provenance: textbook level only, Parr / Pezzulo / Friston, Active Inference, MIT Press 2022. Preprint Polzin et al. 2026 (Zenodo DOI 10.5281/zenodo.19785799, MIT) is the mathematical foundation only, fenced UNREFEREED (Layer-1 AI-executable audit complete; Layer-2 human expert review PENDING).
sha256 f3158b737ac41b7f — of the original file, so what was ingested stays checkable.
Plain — written for this website, not the source document
One engine, reused. That is the whole of the claim made here, and the chapter is careful to keep it that size. The program around it is a simulation, a bounded peek at a toy world and never a person. So showing the same engine at work in a toy cell, in a cardio-renal model and in precision-weighting is a fact about reuse, not about any scale being finished. It is checked mechanically. Each public lab keeps its page engine pinned to a single master copy by a parity test that fails the build if the two ever drift apart. New labs are built by copying a working one and swapping only the observation or generative model, which is the chapter's recurring teaching point: the observation model is not the engine. Every scale is then recorded at its own gate, with the negative that bounds it printed alongside.
Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is f3158b737ac41b7f
Clear — written for this website, not the source document
The organising claim here is architectural and checkable: one engine, reused, across scales. The chapter is careful that this buys reuse and parity, and nothing about whether any scale is finished. Everything the engine runs is a simulation — a toy world, and not a person.
Each public lab is a self-contained page with an inline engine, mirrored by the same physics written out once in a single master copy. The two are pinned together by a parity test that fails the build if they diverge. New labs are made by duplication: copy a working lab, swap only the observation or generative model, and keep the variable names identical so the downstream planner code is reused as written. The discipline carries no empirical result that could overturn it. The only thing that would show it broken is a governance failure, such as a parity test that does not fail the build when the two engines drift. A companion guard fails the build if framing copy, a citation, accessibility or labelling regresses.
Three scales are then recorded, each at its own gate and each with a travelling negative. At the cellular scale a benchmark whose bar was set before the build has the controller topping most modes while honestly losing three named ones, and citing the win without the losses is an overclaim. At the physiological scale a cardio-renal teaching lab re-expresses a published clinical model in inference language, framing a heart attack as a failure of the prediction loop. Its limit is built into the artifact rather than left in the prose: it is not a clinical tool and not a diagnostic instrument, and the resemblance is an interpretive act rather than a measurement. At the cognitive scale three precision knobs sweep a map into distinct behavioural regimes. That work is recorded as test-covered and parity-checked, which is a faithful demonstration of the engine, not a capability shown on data set aside for the test.
One referenced source gets the hardest limit of all. A workbench that is architecturally the most substantive design seed carries no recorded execution evidence at all, so its patterns are a charter and must never be presented as results.
The lead test is mechanical, and it runs two ways. If a parity test does not fail the build when a page's engine and the master copy drift apart, the organising claim is false. It is false too if a lab duplication turns out to have needed engine changes rather than only a model swap.
Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is f3158b737ac41b7f