UNI Universal Natural Intelligence

Wiki · The Encyclopedia

Mv2 - Open falsification (the public invitation)

The Encyclopedia · encyclopedia/wing-M/Mv2-open-falsification.md @ 575fc93d9d31 (main) — opens the published snapshot e850f872196d

How to read this page

Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.

The Encyclopedia is the UNI method written out as a reference work: 39 pages, arranged in wings, setting out what the programme is attempting and why it is built the way it is. This is where the ideas are explained in order and in prose, rather than as code, as runbooks, or as dated receipts.

Every chapter is authored against two ledgers and never ahead of them. One records what UNI has built, and the evidence class of each claim. The other records nature's own regularities, kept separate on purpose. That way a fact about biology is never quietly reused as a fact about the software. Where a chapter and a ledger disagree, the chapter is the thing that is wrong. Every chapter closes with an invitation to falsify it, and a recorded negative is published beside the result it qualifies rather than after it.

Read "How to read this work" first. It is the evidence constitution: the classes, the four ledger states, and the rule that a finished chapter is not the same as a working system. Then the calibration ledger, which carries the figures every other chapter is required to use.

What it is not: a description of a person or of a mind. The programme calls itself a developmental active-inference simulation, a bounded peek into a toy world, and its own index prints how much of the developmental ladder has actually been earned — roughly two rungs out of eleven or more. It is also not a report of what is running today. For what ran, and when, go to the evidence record.

Your browser cannot switch reading levels, so the document itself is shown.

Precise — the source document

This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.

The UNI program does not ask you to trust it. It asks you to try to break it. That is the whole of this chapter: an invitation, addressed to any reader with the patience to re-run a benchmark, to falsify the published claims. The invitation is meaningful only because it is paid for in advance, in the currency that matters most to a science-honest reference work: the program publishes, at the top of its own public leaderboard, the modes where it loses. An invitation to falsify a result that hides its failures is theatre. An invitation that ships the failures first is a wager. This chapter is the wager.

The honest position is fixed and is not softened anywhere below. UNI is a developmental active-inference SIMULATION, a bounded peek inside a toy world, never a person and never a mind. Roughly two of eleven-plus developmental rungs are earned. The program tops SOME registered modes and loses NAMED others. Nothing here claims it has won the field, and no significance verdict in this work is ever decided on one seed or on a point estimate.

The losing leaderboard (the load-bearing artifact)

The flagship demonstration of the invitation is the Cell Lab, an open, pre-registered falsification benchmark (ledger row L1.1, Class C, dev-gate / held-out). A hidden 216-state "service cell" (factor sizes [4,3,3,3,2], ten actions) is perturbed by disturbance families it never announces, and observation-only controllers compete to keep it inside a viable set: UNI active inference, a rule-based SRE heuristic, a random baseline, a small neural-net (MLP) baseline, and, through a "Prove UNI Wrong" upload slot, an outside structural controller. UNI tops the leaderboard on most modes: over six seeds it beats random 7/7 (significant in six), the rule-based controller 6/7, and the neural baseline 5/7.

That PASS is published only with the negative that travels beside it, and the two must never be separated. In the same pre-registered benchmark UNI honestly LOSES on three named modes (ledger row L1.2, Class C, NEGATIVE): on database_flaky the rule-based SRE wins (0.803 vs 0.759); on memory_leak the neural net wins (0.810 vs 0.740); on cpu_noisy_neighbor the neural net wins (0.824 vs 0.749), and on that last mode the UNI-versus-random margin is not even statistically significant. These losses are recorded in FALSIFICATION.md and shown at the TOP of the live leaderboard, not buried in a footnote. The honest reading of L1.1 + L1.2 together is the one the ledger fixes: UNI is "good but not sovereign." It is a strong observation-only controller that is genuinely beaten on three specific disturbance families by simpler methods, and it says so first.

The falsifier is operable and public. The recorded losses are themselves falsifiable claims: if, on the pre-registered benchmark, UNI's RecoveryScore CI separates ABOVE baseline on any of those three modes, the recorded loss has failed to replicate and the row is overturned. Symmetrically, the win is falsified the moment its RecoveryScore CI fails to separate from controls where a win is claimed, or a honesty fence fails.

Why the invitation is sincere, not decorative

Sincerity here is mechanical, not rhetorical. Three disciplines convert "we welcome scrutiny" from a slogan into something a build can fail on.

First, honesty-as-a-test (ledger row M29, method). The benchmark carries eight explicit honesty "fences" and a cell_framing_guard.ts test that FAILS THE BUILD if the page's framing copy, its citation of the preprint DOI, its accessibility, or its "does not reproduce Rao's method" labeling regresses. A claim that drifts upward cannot quietly ship; it breaks the test suite. The falsifier for the discipline itself is precise: if framing, DOI, or labeling regresses WITHOUT the framing_guard failing, the guard is not doing its job.

Second, the statistics gate, not vibes (also M29). Significance is decided by a bootstrap 95% CI on the median paired difference, which must exclude zero, computed with a seeded PRNG (Mulberry32) and a committed result cache, never from a single seed and never read off a point estimate. This is the operable form of the program's standing verdict rule: a capability verdict is the CI bound that excludes the threshold, never the point estimate. It is also the lead falsifier of this entire chapter. If any significance verdict in this work is shown to have been decided on one seed or on a point estimate, the verdict is void.

Third, No-Exit Discipline (ledger row M6, method). A negative is a result, not an exit. The only legitimate place to rest is a proven working solution OR a published, exhausted, falsifiable bound, after which the honest move is to redirect to the axis where the system genuinely excels, never to go silent on an open gate. The discipline is itself falsifiable: exiting on a partial or a negative WITHOUT a published exhausted bound, or treating an undischarged sign-to-park as closed, is the violation.

These disciplines are the same ones the public calibration ledger enforces (chapter E1). That ledger carries the recorded snapshot of 882 rows (350 PASS / 0 FAIL / 183 NEGATIVE / 349 PENDING) as a provenance-flagged snapshot, not as a re-countable headline (the 183-negative figure is not independently re-derivable from the carded rows, so it is cited as a snapshot and the negatives the ledger body actually enumerates are the ones featured). It also carries the down-only calibration record: a peer-reviewed honesty audit (os-cycles 39 to 53) moved six headline claims DOWN to their measured values, "cleared on TWO boxes" to ONE box, "the mind survives a patch" to infrastructure continuity only, "5 stacks" to "4 distinct stacks / 5 runs." The overstatements lived in headlines, not in fabrication, and the calibrated figures are the ones carried forward. Calibration moves down only; the fence gets louder under pressure, not wider.

What is NOT claimed in Mv2

  • Ceiling: That UNI has "won the field," that an open invitation to falsify is itself proof the science is settled, or that publishing losses makes the wins larger, is NOT shown. The most we claim is exactly this: on the pre-registered Cell Lab benchmark UNI tops most modes (beats random 7/7, rule-based 6/7, neural 5/7) and is honestly beaten on three NAMED modes (database_flaky, memory_leak, cpu_noisy_neighbor), with every significance verdict gated on a bootstrap CI excluding zero across multiple seeds, never on one seed or a point estimate.
  • Fences engaged: Red line 7 (never raise a claim above its source evidence class: L1.1/L1.2 are Class C, M6/M29 are method, and none is carded higher). Red line 12 and the vocabulary-leak guard (this is a Movement chapter, but it cites no internal channel handle and features no patent-level math). The standing program framing (developmental SIMULATION, toy world, ~2 of 11+ rungs) is printed, not softened. The forbidden-phrasings list (SIGNED, 2026-06-27) is honored: no "beat LLMs," no "won the field," no awareness claim appears.
  • Negatives that travel with this claim (cite alongside, never strip): L1.2 (the three named Cell Lab losses, top-of-leaderboard) MUST appear next to L1.1 (citing the leaderboard win without the three losses is an overclaim and fails review). The E1 down-only calibration record (the six headlines moved DOWN) travels with any appeal to the program's honesty as evidence of sincerity.
  • Parked / owed: No sign-to-park is owed by this chapter; it asserts a posture and two published rows, not a parked frontier. It inherits the program-wide owed item only by reference: the L7.6 / T2 frontier sign-to-park (UNI_CONSULT_5) remains owner-relayed and not yet captured, and is not discharged by writing about open falsification.
  • One-line honest summary a skeptic could not dispute: A program that prints its own three named losses at the top of a pre-registered, seed-bootstrapped leaderboard, and gates every verdict on a CI that excludes zero, has earned the right to invite falsification, and has claimed nothing beyond being good but not sovereign.

Falsify this

Take the Cell Lab benchmark as published. Re-run it under its committed seeds and the Mulberry32 PRNG, and recompute the bootstrap 95% CI on the median paired difference for database_flaky, memory_leak, and cpu_noisy_neighbor. If UNI's RecoveryScore CI separates ABOVE the winning baseline on any of those three modes, the recorded loss has not replicated and L1.2 is overturned. Equally, if any significance verdict in this benchmark is shown to rest on a single seed or a point estimate rather than a CI that excludes zero, the verdict fails the program's own statistics gate (M29) and is void. The invitation is real precisely because both of these tests are ones you can run, and ones the program has wagered it will survive.

Sources

  • Ledger rows L1.1 (Cell Lab PASS, Class C) and L1.2 (the three named NEGATIVE losses, Class C); method rows M29 (honesty-as-a-test + the statistics gate) and M6 (No-Exit Discipline); chapter E1 (the public calibration ledger; the os-cycles 39 to 53 down-only audit; the provenance-flagged 882/183 snapshot). Source of truth: ../CLAIM-LEDGER.md.
  • Authoring spec: ../MASTER-PLAN.md, section Mv2, with the PART I Evidence Constitution (FM-1), the A-U rubric (FM-2), the red lines (FM-3), and the "what is NOT claimed" template (FM-4).
  • Narrative grounding (PII-redacted digests): ../../curated/uni-precision-digest.md (the Cell Lab falsification benchmark, the eight fences, cell_framing_guard.ts, the leaderboard with losses on top, the bootstrap-CI statistics gate, the recorded NEGATIVE losses); ../../curated/uni-gpt-digest.md (No-Exit Discipline, the append-only ledger, the calibration-down audit). Archive pointers: …-Universal-Natural-Intelligence-Precision and …-UNI-GPT (local-only, never committed; no PII).

sha256 e571e102ed33bae3 — of the original file, so what was ingested stays checkable.

Plain — written for this website, not the source document

Written for this website — not the document. This is a plain-language retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

Do not trust the program, try to break it. That is the invitation the chapter makes, and what it offers up for breaking is a simulation — a toy world, not a person. What makes the invitation worth anything is that it is paid for in advance. The program publishes, at the top of its own public leaderboard, the modes where it loses. An invitation to falsify a result that hides its failures is theatre; one that ships the failures first is a wager. The clearest example is an open benchmark whose rules were written down before anything was run, in which a hidden toy cell is disturbed and several controllers compete to keep it viable. The program's controller tops most modes and honestly loses three named ones, and on one of those it cannot be separated from chance. The calibrated phrase the chapter insists on is that this is good but not sovereign.

Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is e571e102ed33bae3

Clear — written for this website, not the source document

Written for this website — not the document. This is a clearer retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

The chapter states one posture and backs it with two published rows in the ledger, a record added to and never edited, plus the disciplines that make the posture checkable. What it hands to a critic is a simulation — a toy world, and not a person.

The benchmark at its centre is open, and its rules were registered before anything was built. A hidden toy service cell is perturbed by disturbance families it never announces, and observation-only controllers compete to keep it inside a viable set. Those controllers are an active-inference controller, a rule-based heuristic, a random baseline, a small neural network, and an outside entry uploaded through a channel that invites people to prove the program wrong. The program's controller tops the leaderboard on most modes.

That result is never published alone. In the same benchmark it loses on three named modes, and those losses are shown at the top of the live leaderboard rather than in a footnote. On one of them the margin over random is not significant. The honest reading the ledger fixes is good but not sovereign: a strong observation-only controller that can name the exact families where a simpler rule or a small network does better.

Three disciplines turn welcoming scrutiny into something a build can fail on. A framing guard test fails the build if the page's copy, its citation, its accessibility or its labelling regresses, so a claim that drifts upward breaks the suite. A statistics gate decides significance by a bootstrap interval on the median paired difference that must exclude zero, under a seeded generator with a committed cache, never on one seed and never off a point estimate. And a no-exit discipline says the only legitimate rest is a working solution or a published, exhausted, falsifiable bound.

The chapter also carries the program's down-only calibration record, in which an audit pulled six headline claims back to their measured values, and it marks the recorded snapshot of negatives as a snapshot rather than a re-countable headline.

The result that would overturn all of it is one a reader can run. Re-run the benchmark under its committed seeds and recompute the intervals for the three losing modes. If the program's score separates above the winning baseline on any of them, the recorded loss has not replicated and the row is overturned.

Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is e571e102ed33bae3