UNI Universal Natural Intelligence

Wiki · The Encyclopedia

S-L1 - Cellular: the Cell Lab (with its honest losses)

The Encyclopedia · encyclopedia/wing-S/S-L1-cellular.md @ 575fc93d9d31 (main) — opens the published snapshot e850f872196d

How to read this page

Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.

The Encyclopedia is the UNI method written out as a reference work: 39 pages, arranged in wings, setting out what the programme is attempting and why it is built the way it is. This is where the ideas are explained in order and in prose, rather than as code, as runbooks, or as dated receipts.

Every chapter is authored against two ledgers and never ahead of them. One records what UNI has built, and the evidence class of each claim. The other records nature's own regularities, kept separate on purpose. That way a fact about biology is never quietly reused as a fact about the software. Where a chapter and a ledger disagree, the chapter is the thing that is wrong. Every chapter closes with an invitation to falsify it, and a recorded negative is published beside the result it qualifies rather than after it.

Read "How to read this work" first. It is the evidence constitution: the classes, the four ledger states, and the rule that a finished chapter is not the same as a working system. Then the calibration ledger, which carries the figures every other chapter is required to use.

What it is not: a description of a person or of a mind. The programme calls itself a developmental active-inference simulation, a bounded peek into a toy world, and its own index prints how much of the developmental ladder has actually been earned — roughly two rungs out of eleven or more. It is also not a report of what is running today. For what ran, and when, go to the evidence record.

Your browser cannot switch reading levels, so the document itself is shown.

Precise — the source document

This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.

UNI is a developmental active-inference SIMULATION: a no-backprop, nested-Markov-blanket model of a developing organism, never a person, never a mind. This chapter covers the first rung above the genome on the HUMAN-HGM-001 ladder, the cellular level, where the question is whether a controller can keep a tiny simulated cell inside its viable set when an unannounced disturbance tries to push it out. The honest headline here is deliberately modest, and it is printed exactly as the ledger frames it: UNI is good but not sovereign. On the Cell Lab benchmark UNI tops the leaderboard on most modes, and it honestly loses on three named modes, with those three losses shown at the top of the live leaderboard rather than hidden. That pairing, a benchmark win delimited by its own published losses, is the whole point of this chapter. Honest program position, printed and never softened: about 2 of 11+ developmental rungs earned.

What the Cell Lab is

The Cell Lab is an open, pre-registered falsification benchmark, honest by construction rather than by promise. A hidden 216-state service cell (a discrete world built from factor sizes [4, 3, 3, 3, 2], with 10 available actions) is perturbed by families of disturbances it never announces in advance. Observation-only controllers then compete to keep the cell inside its viable set: a UNI active-inference controller, a rule-based SRE heuristic, a random baseline, a small neural-net (MLP) baseline, and, via a "Prove UNI Wrong" upload path, a Rao-style structural controller and a UNI translation of it.

Three design choices make the benchmark a genuine falsification instrument rather than a demo:

  • The scoring rule is fixed in advance. Controllers are scored by RecoveryScore, and significance is decided not by a single seed but by a bootstrap 95% confidence interval on the median paired difference, which must exclude 0. The verdict is the CI bound that excludes the threshold, never a point estimate, and never a lucky seed.

  • The losses are surfaced, not buried. The leaderboard is built to display UNI's losses at the top. The honest disconfirmation table lives in FALSIFICATION.md alongside the claims file, so a skeptic reads the failures before the wins.

  • The framing is test-enforced. Eight explicit honesty fences plus a framing_guard test (cell_framing_guard.ts) fail the build if the page's framing copy, the cited DOI, the accessibility labeling, or the "this is our own schema, not a reproduction of Rao's verified method" labeling ever regresses. Honesty is wired as a passing test, not left to a reviewer's vigilance.

The challenge-upload schema is a structural knowledge graph framed after Mikkilineni 2022 (DOI 10.3390/info13010024), and it is explicitly labeled as the program's own schema, not a reproduction of Rao's method. That label is one of the things the framing_guard protects.

The PASS, with its losses in the same passage (L1.1 and L1.2)

The positive claim and its paired negative are recorded as two ledger rows that must always be read together. Citing the win without the losses is an overclaim and fails review.

  • L1.1 [Class C, dev-gate / held-out, proven]. The Cell Lab benchmark is built and run as described above (216-state service cell, observation-only controllers, RecoveryScore, bootstrap 95% CI, 8 honesty fences plus the framing_guard test), and UNI tops the leaderboard on most modes. Across the recorded run (6 seeds, 80 ticks, depth-2 planning) UNI beats the random baseline 7 of 7 modes (significant in 6 of them), the rule-based SRE heuristic 6 of 7, and the neural-net baseline 5 of 7. The win is real and it is held to a CI bound, which is exactly why the three exceptions matter.

  • L1.2 [Class C, negative]. UNI honestly loses on three named modes, and these are top-of-leaderboard published content:

    • database_flaky: the rule-based SRE controller wins, 0.803 vs 0.759.
    • memory_leak: the neural-net controller wins, 0.810 vs 0.740.
    • cpu_noisy_neighbor: the neural-net controller wins, 0.824 vs 0.749, and here the result is sharper still: even the UNI-versus-random margin is not significant. On this one mode the active-inference controller cannot be separated from chance.

These three losses are recorded in FALSIFICATION.md and shown at the top of the live leaderboard. They are the deliverable, not an embarrassment to be footnoted. They are precisely what licenses the calibrated phrase good but not sovereign: a controller that wins most modes against tuned and learned baselines, and that can name the exact modes where a simpler rule or a small neural net does better. A reader who only saw the wins would carry away a false picture; the program's own benchmark refuses to let that happen.

The exact ceiling on this rung

The single calibrated claim, and nothing stronger: under a pre-registered, observation-only falsification benchmark, UNI's active-inference controller keeps a 216-state toy service cell inside its viable set better than tuned and learned baselines on most modes, and worse on three named modes, with every verdict decided by a bootstrap CI that excludes 0. That is the cellular-viability and homeostasis result, at the cellular and zygote end of the ladder only.

What this rung emphatically is not: it is not "sovereign" control, not autonomy, not life, not a cell that maintains itself in the world. It is a controller scored on a hidden toy world, in a toy benchmark, with its losses published. "Active inference" is the framing lens for the controller, not a demonstrated general capability, and the win is a benchmark result on a 216-state simulation, never evidence of comprehension, awareness, or anything climbing toward a mind.

What is NOT claimed in S-L1 (Cellular: the Cell Lab)

  • Ceiling: That UNI is a self-maintaining or "sovereign" digital cell, or that topping a homeostasis leaderboard is autopoiesis or life, is NOT shown. The most we claim is L1.1 exactly: on the pre-registered, observation-only Cell Lab benchmark UNI's RecoveryScore tops most modes (beats random 7/7, rule-based 6/7, neural 5/7) under bootstrap-CI verdicts, while honestly losing on database_flaky (0.759 vs 0.803), memory_leak (0.740 vs 0.810), and cpu_noisy_neighbor (0.749 vs 0.824, UNI-vs-random not significant). The calibrated phrase is good but not sovereign.
  • Fences engaged: No "created life" / "digital life" / "measurable awareness" as a claim (red line 4); no "active inference demonstrated" (red line 3, "active inference" is the lens for the controller, not a demonstrated loop); no claim raised above its source class, which is Class C here (red line 7); calibration moves DOWN only, so "tops most modes" never becomes "wins" and never becomes "sovereign" (FM-1). This is a toy world, not the real world.
  • Negatives that travel with this claim (cite alongside, never strip): L1.2 in full, the three named losses (database_flaky 0.759 vs 0.803, memory_leak 0.740 vs 0.810, cpu_noisy_neighbor 0.749 vs 0.824 with the UNI-vs-random margin not significant). The PASS (L1.1) may never be published without these three losses in the same view, exactly as the live leaderboard shows them.
  • Parked / owed: Nothing is parked at this rung and no sign-to-park is owed; the negatives are already recorded and published. The standing owed obligation is fidelity: keep the losses visible and let the framing_guard test, not a human's good intentions, enforce the framing.
  • One-line honest summary a skeptic could not dispute: On an open, pre-registered, observation-only benchmark with CI-gated verdicts, UNI's controller wins most homeostasis modes against tuned and learned baselines and loses three named ones, and it publishes those losses at the top of its own leaderboard.

Falsify this

The lead falsifier is operable and pre-registered. Re-run the Cell Lab benchmark under its registered protocol; the L1.1 win is falsified if UNI's RecoveryScore CI fails to separate from the controls on any mode where a win is claimed, or if a fence or the framing_guard test fails (an overclaim is emitted), or if the bootstrap CI is shown to be miscomputed. The L1.2 losses carry their own mirror-image falsifier: on the pre-registered benchmark, UNI's RecoveryScore CI separates above baseline on database_flaky, memory_leak, or cpu_noisy_neighbor (the recorded loss fails to replicate). Either outcome would move a ledger row; neither has.

Sources

  • Digest: curated/uni-precision-digest.md (the Cell Lab falsification benchmark: 216-state service cell, RecoveryScore, bootstrap 95% CI, the 8 honesty fences and framing_guard, and the three recorded losses at the top of the leaderboard).
  • Digest: curated/uni-mind-digest.md (cellular / early-homeostasis end of the developmental ladder; the negatives-are-content discipline; the no-backprop conjugate-Dirichlet engine reused verbatim).
  • Ledger: encyclopedia/CLAIM-LEDGER.md, rows L1.1 (PASS, Class C) and L1.2 (NEGATIVE, Class C), section "L1 - Cellular".
  • Archive pointers (not read here; PII-fenced): the UNI Precision Lab archive (...-Precision) and the uni-mind archive (...-uni-mind).

sha256 e0ee5adbddd74f7a — of the original file, so what was ingested stays checkable.

Plain — written for this website, not the source document

Written for this website — not the document. This is a plain-language retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

A controller is dropped into a small simulated cell and asked to hold it inside a viable range while disturbances it was never told about push against it. That is the cellular rung of the developmental ladder, and the headline fixed for it in the ledger, a list added to and never edited, is deliberately small: good but not sovereign. On an open benchmark whose rules were fixed before the run, the program's controller leads on most disturbance modes and is beaten on three named ones. Those three defeats sit at the top of the live leaderboard rather than being tucked out of sight. The pairing is the whole point of the chapter, a win published together with the exact conditions under which it does not hold. Nothing here is autonomy and nothing here is life. What is on offer is a scoring result on a hidden toy world, with the failures shown first.

Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is e0ee5adbddd74f7a

Clear — written for this website, not the source document

Written for this website — not the document. This is a clearer retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

The benchmark is built to be a falsification instrument rather than a demo. Everything it scores happens inside a simulation — a toy world, never a person. A hidden discrete service cell is perturbed by families of disturbances it never announces in advance, and observation-only controllers compete to keep it inside a viable set. Those controllers are an active-inference one, a rule-based heuristic, a random baseline, a small neural network, and outside entries uploaded through a channel that invites a challenger.

Three design choices carry that. The scoring rule is fixed in advance, and significance is decided by a bootstrap interval on the median paired difference that must exclude zero, never by a lucky seed and never off a point estimate. The losses are surfaced rather than buried, since the leaderboard is built to display them at the top and a separate falsification file carries the disconfirmation table beside the claims. And the framing is test-enforced. Honesty limits written into the page, plus a guard test, fail the build if the page's copy regresses. The same goes for its citation, its accessibility labelling, and its statement that this is the program's own schema rather than a reproduction of somebody else's method.

Two rows in the ledger, which is added to and never edited, must always be read together. The positive row records that the benchmark is built and run as described and that the controller tops most modes, beating the random, rule-based and neural baselines on most of them. The negative row records three named modes where it loses, and on one of those the margin over random is not even significant, so on that mode the controller cannot be separated from chance.

The chapter is blunt that citing the win without the losses is an overclaim that fails review. The calibrated claim is exactly this. Under a benchmark whose rules were fixed before the run, with every controller only watching and never told what is coming, the controller keeps a toy service cell inside its viable set. It does that better than tuned and learned baselines on most modes, and worse on three named ones. Every verdict is decided by an interval that excludes zero.

What this step is not is stated as carefully: not sovereign control, not autonomy, not life, not a cell that maintains itself in the world. What would overturn the chapter runs both ways, since the win falls if the interval fails to separate where a win is claimed, and each recorded loss falls if it fails to replicate.

Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is e0ee5adbddd74f7a