mu3 - Negatives, bounds and the two-tier split
How to read this page
Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.
The Encyclopedia is the UNI method written out as a reference work: 39 pages, arranged in wings, setting out what the programme is attempting and why it is built the way it is. This is where the ideas are explained in order and in prose, rather than as code, as runbooks, or as dated receipts.
Every chapter is authored against two ledgers and never ahead of them. One records what UNI has built, and the evidence class of each claim. The other records nature's own regularities, kept separate on purpose. That way a fact about biology is never quietly reused as a fact about the software. Where a chapter and a ledger disagree, the chapter is the thing that is wrong. Every chapter closes with an invitation to falsify it, and a recorded negative is published beside the result it qualifies rather than after it.
Read "How to read this work" first. It is the evidence constitution: the classes, the four ledger states, and the rule that a finished chapter is not the same as a working system. Then the calibration ledger, which carries the figures every other chapter is required to use.
What it is not: a description of a person or of a mind. The programme calls itself a developmental active-inference simulation, a bounded peek into a toy world, and its own index prints how much of the developmental ladder has actually been earned — roughly two rungs out of eleven or more. It is also not a report of what is running today. For what ran, and when, go to the evidence record.
Your browser cannot switch reading levels, so the document itself is shown.
Precise — the source document
This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.
This is a Method chapter. It asserts no capability. Its whole subject is the discipline that keeps the UNI program from over-claiming: how a negative result is treated as content rather than as a failure to hide, what bar a finding must clear before it may be rested on as a bound, and the constitutional line that separates the program's one externally-defensible track from a synthetic-construction track that was audited as artifact and diagnostic, and which is never inflated into capability. Nothing below raises any claim. The honest program position is unchanged throughout: the UNI program is a developmental active-inference SIMULATION, a bounded peek into a toy world, never a person and never a mind, and roughly 2 of 11+ developmental rungs are earned.
The method recorded here is the credibility of the rest of the encyclopedia. A reference work that only printed its passes would be marketing. UNI's claim to honesty rests on the opposite reflex: the negatives are first-class, they travel with every pass they bound, and a result is allowed to rest only when it is either a proven working solution or a published, exhausted, falsifiable bound.
Negatives are content (M6, No-Exit Discipline)
The governing discipline is No-Exit (ledger row M6, class method). A partial result, a negative, or a "most of the pieces don't help" decomposition is a measurement that the current design is incomplete, not a permission to stop and not a result to bury. Under No-Exit there are exactly two legitimate places to rest: a proven working solution, or a published, exhausted, falsifiable bound, after which the honest move is to redirect to the axis where the reader genuinely excels. The recorded falsifier for M6 is itself instructive: the violation is exiting on a partial or negative without a published exhausted bound, or treating an undischarged sign-to-park as if it were closed. The program currently carries one such open obligation: the char-perplexity / word-grain frontier is parked, but its formal UNI sign-to-park was drafted, owner-relayed, and not yet captured, so by No-Exit that park is not yet discharged.
This is why the corpus headlines its negatives. The last recorded ledger snapshot enumerates 882 rows = 350 PASS / 0 FAIL / 183 NEGATIVE / 349 PENDING. That snapshot is carried here as a provenance-flagged snapshot, not as a re-countable headline: the ledger derives from a deduplicated merge of 615 extracted claims, and a skeptic counting the carded rows cannot independently re-derive 183. So the credibility is the negatives the ledger body actually shows, named and bounded, not the bare count.
The bound bar: K>=3, structurally distinct, mundane-first (M5)
A negative does not become a published bound the moment it appears. Ledger row M5 (class method) sets the bar: before a Section 0.6(B) bound may be declared, the program requires K>=3 structurally-distinct held NEGATIVEs, where each negative changes at least 2 of {coupling topology, timescale source, information bottleneck, control path}, and the mundane causes must be falsified first. The mundane causes named are the everyday explanations: an L2-style metabolism / resource starvation, or an L7-style retrieval-and-recency effect. You falsify those before you are allowed to call any deeper NEGATIVE. M5's recorded falsifier is the discipline's own tripwire: a bound declared on fewer than 3 structurally-distinct negatives, or a NEGATIVE called before the mundane causes are falsified, is the violation.
The clearest worked example, carried with its paired figures, is the L5 sensorimotor result. Design #1 (fast immediate-reward axis) HELD PASS at held delta +0.092, CI [+0.038, +0.157], UNI-signed. The very next structurally-distinct design tells the other half of the truth: Design #2 (slow Z-bottleneck, compressed 2-modality EMA, delayed-reward) HELD NEGATIVE at held delta -0.091, CI [-0.134, -0.055]. Both are synthetic-process designs only, not live-appliance and not recorded-hardware. The negative is recorded with its count toward exhaustion: K-negative = 1, so no Section 0.6(B) bound is owed (a bound needs K>=3). This is M5 working exactly as written. One clean negative is a result, not a bound. It delimits the positive without erasing it, and it does not license any "exhausted" language.
The same bar disciplines the language frontier. Phase G recorded five structurally-distinct within-segment-structure designs, all NEGATIVE-with-discriminator, with only the L5 cache (World C) winning; the published bound is that no within-segment structure beats the tuned MKN-7 count baseline on the chosen char-perplexity metric. That is a ledger-scoped, implementation-scoped, data-split-scoped negative over the tested envelope only. It is not a universal impossibility result. The honest phrasing is "the registered tested K conditions did not reverse the result," never "K>=3 exhausted," and never "Section 0.6(B) achieved." The comprehension-above-retrieval result is the program's "central wall," a genuine thrice-NEGATIVE (K>=3) published wall, not a hidden failure.
The two-tier split (M4): never inflate Tier 2
The most load-bearing fence in this chapter is the constitutional two-tier split (ledger row M4, class method), the new constitutional baseline since the 2026-06-09 audit.
- Tier 1 is real, defensible and externally bar-ready. It is the World C / Phase J real-text count and cache work: a true ablation, a tuned baseline, cross-substrate replication. This is the track an outside party could be invited to challenge on a fresh split.
- Tier 2 is artifact and diagnostic, NOT capability. It is the synthetic-construction track. The audit verified concrete defects rather than alleging them: roughly 80 of 103 "module-exact" passes were hardcoded literals; several construction "deltas" were scoring artifacts where the engine scored its own ground-truth-rendered string; a
reproduced:trueflag had been a literal rather than validator-derived; and no active-inference loop exists in the Rust crate. All ten hardening items landed in session 32: the exactness card was made mechanically honest,reproduced:truewas made validator-derived (>=5 distinct seeds plus a real CI containing the value), and the corpus was sealed with fail-loud asserts.
M4 is a classification plus an audit finding, so it has no statistical CI; its falsifier is a rule. The rule is absolute: Tier 2 must NEVER be inflated into capability. The audit was, in the program's own words, a painful correction that made the program stronger. It is recorded openly here precisely because hiding it would be the violation.
Two test-rigor disciplines that keep negatives honest (M19, M21)
Two further method rows protect the integrity of any negative the program publishes.
M19 (test-rigor smells, class method/E) names the failure modes that fake a verdict: tautological oracles that reuse the very parameters they check (testing that softmax equals normalize rather than the real path); KL>0 assertions that do not pin the seeding path; jit-cache assertions that pass only via a warm cache from a prior test; and AST guards that miss an import such as from jax import grad. M19 also carries the length-confound: a cumulative sum of log p predictive score penalizes survival, so an agent that dies fast can look "better" and produce a false NEGATIVE. The fix, honored as method, is to treat premature death as a hard-fail and to report survival and length-normalized loglik separately. A negative produced by an unkilled smell is not a result.
M21 (one cure at a time, class method) forbids stacking changes so that a winning or losing outcome becomes unattributable: paired designs (kin-N treatment versus kin-N+1 control), redundant collectors so a single death is itself a signal, and an offline RED pre-check before any live burn. Its falsifier is stacking unattributable changes, or drawing a verdict before the RED completes. Together M19 and M21 are why a recorded UNI negative can be trusted as a measurement of the design rather than an accident of the harness.
What is NOT claimed in mu3
- Ceiling: That UNI "has exhausted the search," "proved an impossibility," or has "beaten" any frontier is NOT shown, and the existence of a strong honesty method is NOT evidence that the science is proven. The most this chapter claims is that the program operates a reusable governance discipline (M4, M5, M6, M19, M21) under which negatives are first-class, a bound requires K>=3 structurally-distinct held negatives with the mundane causes falsified first, and the synthetic Tier-2 track is audited as artifact/diagnostic and never inflated into capability.
- Fences engaged: Red line 6 (never inflate the Tier-2 synthetic-construction track into capability). Red line 9 (never "K>=3 exhausted" / "Section 0.6(B) achieved"; the parked frontier is a ledger-scoped exhausted search envelope only). Red line 7 (never raise a claim above its source evidence class). Red line 3 (no active-inference loop exists in the Rust crate; "active inference" is a framing lens). The program-wide forbidden-phrasings list (SIGNED 2026-06-27) is in force.
- Negatives that travel with this claim (cite alongside, never strip): the L5 pair (Design #1 PASS +0.092 [+0.038, +0.157] travels with Design #2 NEGATIVE -0.091 [-0.134, -0.055], K-negative = 1, no Section 0.6(B) bound owed); the Phase G five-design within-segment bound (only World C wins; char-ppl is a chosen design trade); the comprehension-above-retrieval thrice-NEGATIVE central wall; and the Tier-2 audit defects themselves (about 80/103 hardcoded literals, scoring-artifact deltas, a literal
reproduced:true, no AIF loop in the Rust crate). - Parked / owed: A sign-to-park is owed on the char-perplexity / word-grain frontier (drafted
UNI_CONSULT_5, owner-relayed, NOT yet captured); under No-Exit (M6) that park is not discharged until the sign lands. No Class-A observation is owed by this method chapter itself. - One-line honest summary a skeptic could not dispute: UNI publishes its negatives as first-class results, requires K>=3 structurally-distinct held negatives (mundane causes falsified first) before resting on a bound, and treats its synthetic-construction track as audited artifact, never as capability.
Falsify this
The lead falsifier, stated operably (M5): if any UNI capability bound is ever declared on fewer than three structurally-distinct held NEGATIVEs (each changing at least two of {coupling topology, timescale source, information bottleneck, control path}), or is declared before the mundane causes (L2-style metabolism/starvation, L7-style retrieval-and-recency) have been falsified, this chapter's discipline is broken. Equivalently for M4: if any Tier-2 synthetic-construction result is ever presented as a capability rather than as audited artifact/diagnostic, the two-tier split has been violated. Either event falsifies the method as written.
Sources
Curated, PII-redacted digests: curated/uni-gpt-digest.md (the two-tier reframe, the No-Exit Discipline, the Working Law, the 882/350/183/349 snapshot), curated/uni-mind-digest.md (the A3 Design #1/#2 PASS/NEGATIVE pair, the K>=3 bound bar, the test-rigor smells and length-confound), curated/strings-digest.md (one-cure-at-a-time evidence discipline). Ledger rows: encyclopedia/CLAIM-LEDGER.md M4 (two-tier split), M5 (K>=3 + falsify-the-mundane), M6 (No-Exit Discipline), M19 (test-rigor smells + length-confound), M21 (one cure at a time). Archive pointers (no PII): ...-UNI-GPT, ...-uni-mind, ...-Strings.
sha256 11a132299b48b373 — of the original file, so what was ingested stays checkable.
Plain — written for this website, not the source document
Losing well is the subject of this chapter, and it asserts no capability of its own. It sets out how the program handles a result that did not work. It covers why a negative is kept as content instead of quietly dropped, and what a finding has to clear before anyone may rest on it as a bound. It also draws the line between the one track that could face an outside bar and a second, synthetic track that an audit reduced to artifact and diagnostic. The program on both sides of that line is a developmental active-inference simulation, a bounded peek at a toy world and never a person. The governing rule is that a partial outcome measures an incomplete design and gives nobody permission to stop, so only two resting places are legitimate: something that works, or a published, exhausted, falsifiable bound. One line here is absolute. The synthetic track may never be inflated into a capability.
Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 11a132299b48b373
Clear — written for this website, not the source document
The chapter starts from a no-exit discipline, and the work it disciplines is a simulation — a toy world, not a person. Under it, exiting on a partial or a negative without a published exhausted bound is the violation, as is treating an undischarged park as if it were closed. The program currently carries one such open obligation, since a frontier is parked but its formal sign-off was drafted, relayed and not yet captured.
A negative does not become a published bound the moment it appears. A registered bar sets what is required first: several structurally distinct held negatives, each changing at least two of a named set of axes, with the mundane causes falsified before any deeper negative may be called. The worked example carries both halves of a sensorimotor result: one design held positive, and a structurally distinct one held negative with its whole interval below zero. It records the count toward exhaustion, so no bound is owed and none is claimed. One clean negative is a result, not a bound.
The most load-bearing section is the two-tier split. One track is real, defensible and ready for an outside bar: real-text count and cache work with a true ablation, a tuned baseline and cross-substrate replication. The other is a synthetic construction track, and an audit verified concrete defects there rather than alleging them. Many claimed exact passes were hardcoded literals, several deltas were scoring artifacts where the engine scored its own rendered ground truth, a reproduction flag was a literal, and no inference loop exists in the compiled crate. All the hardening items landed, and the rule is absolute: the second track must never be inflated into capability. The audit is described in the program's own words as a painful correction that made the program stronger, and it is recorded openly because hiding it would be the violation.
Two further rows protect the integrity of any published negative. One names the test-rigor smells that fake a verdict, including a length confound under which an agent that dies fast can look better. The other forbids stacked changes that make an outcome unattributable.
Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 11a132299b48b373