mu2 - Bars-before-build, held-once, CI-as-verdict
How to read this page
Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.
The Encyclopedia is the UNI method written out as a reference work: 39 pages, arranged in wings, setting out what the programme is attempting and why it is built the way it is. This is where the ideas are explained in order and in prose, rather than as code, as runbooks, or as dated receipts.
Every chapter is authored against two ledgers and never ahead of them. One records what UNI has built, and the evidence class of each claim. The other records nature's own regularities, kept separate on purpose. That way a fact about biology is never quietly reused as a fact about the software. Where a chapter and a ledger disagree, the chapter is the thing that is wrong. Every chapter closes with an invitation to falsify it, and a recorded negative is published beside the result it qualifies rather than after it.
Read "How to read this work" first. It is the evidence constitution: the classes, the four ledger states, and the rule that a finished chapter is not the same as a working system. Then the calibration ledger, which carries the figures every other chapter is required to use.
What it is not: a description of a person or of a mind. The programme calls itself a developmental active-inference simulation, a bounded peek into a toy world, and its own index prints how much of the developmental ladder has actually been earned — roughly two rungs out of eleven or more. It is also not a report of what is running today. For what ran, and when, go to the evidence record.
Your browser cannot switch reading levels, so the document itself is shown.
Precise — the source document
This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.
This chapter documents a discipline, not a science result. It carries no capability claim and earns no developmental rung. What it describes is the procedure the UNI program uses to decide whether any other claim is allowed to exist: how a bar is registered before a build begins, how a held-out set is touched exactly once, how a verdict is read off a confidence-interval bound rather than a point estimate, and how the word "reproduced" is forced to mean something. The honest program position is unchanged by anything written here: the whole program remains a developmental active-inference SIMULATION, a bounded peek into a toy world, with roughly 2 of 11+ developmental rungs earned. This chapter is part of why those two are believable and why the other nine are still openly marked unearned.
The four method rows below are recorded in the claim ledger as status: method (class method, with E enforcement where the discipline is wired into a server-side board). They are governance patterns, proven and reusable, and they cannot be raised into a capability claim. Where this prose and the ledger disagree, the ledger wins.
Bars before build, held once (M2)
The first rule is ordering. A bar is registered before a build starts: a stated margin against a named threshold, plus a named ablation or control that is expected to collapse the gain if the effect is real. Only then is code written. The held-out set is touched exactly ONCE, behind an atomic seal-before-scoring step and a once-only sentinel; re-scoring the same held set is refused by construction. The run produces a frozen-config artifact and an ordered assert chain, so a verdict cannot be quietly re-derived under a different configuration.
The load-bearing clause is the verdict rule itself: the verdict is the confidence-interval bound that excludes the threshold, never the point estimate. A point estimate that clears a bar is not a pass. The pass is the CI bound. This is what makes the program's one citable empirical PASS legible: World C is recorded as +0.081 nats/char over a tuned MKN-7 count baseline, and the verdict rests on the multi-seed CI [0.0736, 0.0890] (seeds 0-4) whose lower bound, 0.0736, clears the 0.03 bar by 2.4x. The number that decides the verdict is 0.0736, the bound, not 0.081, the estimate. The same shape governs the Phase J specialist reader: held margin +0.105 nats/char with the M-seal CI [0.0975, 0.1128], and the verdict is the lower bound 0.0975.
Every PASS this method certifies travels with its paired negative, because citing a pass without its travelling negative is itself an overclaim that this discipline is meant to catch. World C is a COUNT baseline win, explicitly NOT active inference, NOT comprehension, NOT "talking," NOT "beats LLMs": the program sits roughly 10-15% behind backprop LLMs on char-perplexity by a chosen design trade. Phase J's +0.105 is carried in the ledger beside J.attribution_caveat (NEGATIVE): overall NLL worsens; the gain is a specialist gain, not a general win, and the caveat must always be cited alongside the pass. The very same held-once machinery that signs the passes also signs the symmetric NEGATIVE bound that delimits them: A3 Design #2 (a slow Z-bottleneck motor design) was held NEGATIVE at -0.091, CI [-0.134, -0.055], the upper bound -0.055 sitting below zero. A method that can only ever produce passes is not this method. The negatives are sealed and signed by the identical procedure.
This row's falsifier is operable and blunt: a held set touched more than once, a verdict read off a point estimate instead of the CI bound, or a build started before its bar is registered. Any one of those voids the result the method was supposed to license.
Validator-derived "reproduced:true" (M3)
The word "reproduced" is a load-bearing word, so the program does not let a human or a narrative type it. A reproduced:true flag must be DERIVED by the validator from at least 5 distinct seeds plus a real, non-degenerate confidence interval that contains the value, never written in as a hardcoded literal. This was the central fix of the 2026-06-09 audit, which found a reproduced:true that had been a literal rather than a computed verdict. World C is recorded as the first genuinely validator-derived reproduced:true in the corpus, which is precisely why it is the flagship and the others before it were not.
The fence here is exact and it is in the load-bearing line of this chapter: reproduced:true is validator-derived (at least 5 seeds and a real non-degenerate CI), never a literal. The falsifier is the literal itself: a reproduced:true emitted as a constant, not derived from at least 5 seeds and a real non-degenerate CI, voids the flag.
Contains-baseline plus a load-bearing discriminator (M7)
A pass also has to be the right shape. Every capability claim needs three things at once: (a) a tuned strong baseline (an untuned baseline that the system "beats" proves nothing); (b) a load-bearing discriminator (shuffle-labels, marker-swap, or ablate-to-zero) that COLLAPSES the gain when the mechanism is removed; and (c) a true ablation that is a computed residual, not a hardcoded literal. If removing the mechanism does not remove the gain, the gain was never the mechanism's. World C is measured against a tuned MKN-7, not a strawman; Phase J carries a structure-margin discriminator that is required to collapse the gain under marker-swap, and a contains-KN-OOV control at lambda = 0. The falsifier: a capability claim without a tuned baseline, without a collapsing discriminator, or resting on a hardcoded-literal ablation.
The server-enforced evidence contract (M8)
The discipline above is not left to good intentions; it is wired into a board that refuses to advance without evidence. Every ticket walks a fixed cycle, TODO -> DD -> TDD_RED -> TDD_VERIFY -> TDD_GREEN -> TDD_REFACTOR -> TDD_VALIDATE -> DONE, and each transition is hard-blocked unless a marker-bearing evidence comment is logged first: a RED transition needs a test_/assert/expect marker, a REFACTOR transition needs a refactor/extract marker, a VALIDATE transition logs a full-suite pass. Skipping steps is server-enforced, not a convention. DONE requires at least 2 explicit Y: verdicts per criterion. And a claim linter runs over the ticket comments and auto-downgrades overclaims: in a recorded run it caught the bare word "PROVEN" in a verdict and forced it down to Class E, the held class. This row carries E enforcement because the board actually fails the build when the contract is breached. The falsifier: a phase transition that succeeds without its required marker, a ticket reaching DONE with fewer than 2 Y: verdicts, or the linter failing to downgrade "PROVEN."
A recorded illustration of the contract under pressure: an operator refused to mark a DONE criterion ("verified in heart-lab.html") satisfied because the viewer it referenced belonged to a LATER ticket, and checked in rather than fabricate a Y: verdict. The discipline got louder under urgency, not wider. That is the intended behavior, and it is the whole point of writing the rules into the board rather than into a person's memory.
What is NOT claimed in mu2
- Ceiling: A careless reader might infer that "the UNI program has a rigorous evidence pipeline" means "the UNI science is therefore proven," or that the World C / Phase J figures cited here as method examples are themselves capabilities. NEITHER is shown. The most we claim is that bars-before-build, held-once scoring, CI-as-verdict, validator-derived reproduction, contains-baseline-plus-discriminator, and a server-enforced evidence contract are reusable governance disciplines (class
method,Ewhere board-enforced) that decide whether a claim is admissible. They assert no capability and earn no rung. - Fences engaged: Red line 7 (never raise a claim above its source evidence class) is the spine of this chapter. Red line 5 (never "beats LLMs"; World C is a COUNT baseline, ~10-15% behind backprop LLMs by a chosen trade) and red line 3 (never "active inference demonstrated"; the lens, not a result) fence the example figures. Red line 1 (never AGI / human-level / "understands") and red line 2 (never consciousness / sentience) stand because a method chapter is the place a reader is most tempted to over-generalize.
- Negatives that travel with this claim (cite alongside, never strip): World C +0.081 always with "COUNT baseline, not AIF, not comprehension, not beats-LLMs." Phase J +0.105 always with
J.attribution_caveat(overall NLL worsens; specialist gain only). The held-once machinery is shown signing the A3 Design #2 NEGATIVE (-0.091, CI [-0.134, -0.055]) in the same breath as the A3 Design #1 PASS, to demonstrate that the procedure produces sealed negatives, not only passes. - Parked / owed: Nothing is parked for this chapter itself; it is a discipline, fully recorded as
method. No sign-to-park is owed here and no Class-A observation is owed: a governance pattern is not a runtime capability awaiting a live check. - One-line honest summary a skeptic could not dispute: This chapter describes how the program decides what counts as evidence; it proves no capability, and every example figure is carried with the negative that bounds it.
Falsify this
Lead falsifier, stated operably: show a held-out set in this program touched more than once, or a published verdict read off a point estimate rather than the confidence-interval bound that excludes the threshold, or a reproduced:true that was written as a literal rather than derived by the validator from at least 5 distinct seeds and a real non-degenerate CI. Any single instance falsifies the claim that this discipline is actually in force.
Sources
- Method rows M2 (bars-before-build, held-once, CI-as-verdict), M3 (validator-derived
reproduced:true), M7 (contains-baseline + load-bearing discriminator), M8 (server-enforced DD-TDD evidence contract) -CLAIM-LEDGER.md, section 3. - Authoring spec -
MASTER-PLAN.md, Wing mu, section mu2; front matter FM-1 (Evidence Constitution), FM-2 (A-U rubric), FM-3 (red lines), FM-4 (not-claimed template). - Narrative grounding (PII-redacted):
curated/uni-mind-digest.md(the bars-before-build / ORCHESTRATE loop, the held-once seal + once-only guard, the symmetric A3 PASS/NEGATIVE, calibration-down under pressure) andcurated/uni-precision-digest.md(board-synced TDD, the claim-linter "PROVEN" -> Class E downgrade, the bootstrap-CI statistics gate, the refused DONE criterion under pressure). - Underlying archives (not read here; pointers only):
...-uni-mind,...-UNI-GPT,...-Precision,...-SolutionWright-IdeationExplorer.
sha256 81fd6e1a86b1cfce — of the original file, so what was ingested stays checkable.
Plain — written for this website, not the source document
Admissibility is the subject here: what a claim must survive before the program will let it exist at all. Four method rows carry the chapter, and not one of them is a science result or earns a step on the developmental ladder. The thing they police is a developmental active-inference simulation, a bounded peek at a toy world, so this discipline is most of what makes its few modest claims believable. Ordering comes first. A bar goes on record before any build starts. A sealed set is scored once, and a second scoring is refused by construction. The verdict is then read from the edge of an interval rather than its middle, so clearing a bar with a point estimate does not count. The word reproduced is taken out of a writer's hands too, and made something a validator has to derive rather than type. And the same machinery that signs the program's wins is shown signing a bound that limits them, which is the chapter's own test: a method that can only produce passes is not this one.
Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 81fd6e1a86b1cfce
Clear — written for this website, not the source document
Four method rows carry the chapter, all recorded as governance patterns, with test enforcement where a rule is wired into a board. What the rows police is a simulation — a toy world, and not a person.
The first is ordering. A bar is registered before a build starts, as a stated margin against a named threshold plus a named ablation expected to collapse the gain if the effect is real, and only then is code written. The set held back for scoring is touched exactly once, behind an atomic seal before scoring and a once-only sentinel, so re-scoring is refused by construction. The run produces a frozen configuration and an ordered chain of assertions, so a verdict cannot quietly be re-derived under different settings. Every result the method certifies travels with its paired negative, including the headline reading result, which is a count baseline and explicitly not comprehension, and a specialist language gain carried beside a caveat that the overall score worsens.
The second row makes the word reproduced mean something. The flag must be derived by the validator from several distinct seeds plus a real, non-degenerate interval containing the value, never written in as a literal. That was the central fix of an audit which found exactly such a literal.
The third row says a pass also has to be the right shape. It needs a tuned strong baseline, a discriminator that collapses the gain when the mechanism is removed, and a true ablation that is a computed residual rather than a hardcoded number. If removing the mechanism does not remove the gain, the gain was never the mechanism's.
The fourth row wires the discipline into a board that refuses to advance without evidence. Each transition is blocked unless a marker-bearing evidence comment is logged first, a finished ticket requires explicit verdicts per criterion, and a linter downgrades overclaims automatically. A recorded illustration shows the contract under pressure, where an operator refused to mark a criterion satisfied because the artifact it referenced belonged to a later ticket, and checked in rather than fabricate a verdict.
Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 81fd6e1a86b1cfce