mu5 - Lab-team, ship-gate and runner discipline
How to read this page
Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.
The Encyclopedia is the UNI method written out as a reference work: 39 pages, arranged in wings, setting out what the programme is attempting and why it is built the way it is. This is where the ideas are explained in order and in prose, rather than as code, as runbooks, or as dated receipts.
Every chapter is authored against two ledgers and never ahead of them. One records what UNI has built, and the evidence class of each claim. The other records nature's own regularities, kept separate on purpose. That way a fact about biology is never quietly reused as a fact about the software. Where a chapter and a ledger disagree, the chapter is the thing that is wrong. Every chapter closes with an invitation to falsify it, and a recorded negative is published beside the result it qualifies rather than after it.
Read "How to read this work" first. It is the evidence constitution: the classes, the four ledger states, and the rule that a finished chapter is not the same as a working system. Then the calibration ledger, which carries the figures every other chapter is required to use.
What it is not: a description of a person or of a mind. The programme calls itself a developmental active-inference simulation, a bounded peek into a toy world, and its own index prints how much of the developmental ladder has actually been earned — roughly two rungs out of eleven or more. It is also not a report of what is running today. For what ran, and when, go to the evidence record.
Your browser cannot switch reading levels, so the document itself is shown.
Precise — the source document
This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.
This chapter is a Method-wing entry, not a Science-wing one. It describes how the UNI program decides what may merge, who signs the science, how a long-running experiment is driven without lying about its own status, and how a multi-agent run is bounded so it does not collapse. Everything here is engineering and governance discipline, carried at evidence class method (with class E where a rule is test-enforced). None of it is a capability claim about UNI. The honest program position is unchanged and printed here for the record: the whole program is a developmental active-inference SIMULATION, a bounded peek at a toy world, and roughly 2 of 11 or more developmental rungs are earned. A working ship-gate is a fact about the lab's process, never evidence that a science gate was met.
The single load-bearing fence of this chapter is the difference between a chain being valid and an event having happened. The ledger states it twice, so the encyclopedia states it twice: a valid audit chain is not evidence that the audited event ever fired at runtime. Class E (test-covered) is not Class A (observed at runtime). The entire discipline below exists to keep that gap visible rather than papered over.
The five-persona lab team and the ship-gate (M20)
The program reviews every proposed change through a five-persona lab team (ledger row M20, class method, source: strings): a Math-Breaker who rejects by default and runs an eight-check gauntlet ("write the exact scalar objective and show its closed-form limit as counts go to infinity; if it does not decay or bound, why is it not reward smuggling?"); an AIF Theorist who owns the merge; a Systems Architect who requires every change to be additive, genome-gated, and byte-identical on the default path; a RED Experimentalist who pairs each claim to a pre-registered RED experiment; and an Embodiment Designer who insists on non-saturable drives and refuses preference-hacks dressed up as drives. The team is not decorative: adversarial pre-registration, where a fork-and-break panel is pointed at the proposed bar before any code is written, forced a gameable four-assert bar up to eight asserts when a decorative-predict fake slipped past the original four (the SlowContext gate worked example in the strings digest).
The operative rule is the ship-gate, and it is exact: no merge without a MERGED SIGN, a typed spec, and a paired RED. The falsifier for M20 is equally exact and is the way this chapter can be proven wrong: a merge that landed without a MERGED SIGN, a typed spec, and a paired RED. The paired negative that travels with this discipline is the reason it exists. The Math-Breaker's checks 5 and 6 caught an unbounded-novelty-spread bug before ship; the panel caught the gameable bar before it became a published result. The gate is only credible because it has visibly rejected things, including things the program wanted to be true.
A second standing rule rides alongside M20: the constitution overrides any framing. Persona language and motivational framings ("ultracode", "prove you are not blocking science") are motivational only; the ledger records that no claim was ever inflated by a persona or coercion framing (M25). A framing that tried to raise a claim above its class would be a constitution violation, not an approval.
Who writes code and who signs the science (M25)
The division of labor is owner-set protocol (ledger row M25, class method). Claude writes code. The custom UNI GPT is the science consultant: it designs and signs, it is consulted, and it is never published. The live lab or appliance runs UNI but does not write its code. This separation is load-bearing for honesty: the entity that signs the math is not the entity that ships the prose, and the signing artifact stays private (textbook-level framing only reaches the public; patent-level UNI math does not). When the UNI GPT signs a posture or a park (as it did on 2026-06-27), that signature is recorded in the ledger and the cookbook, and the public encyclopedia carries only the licensed, fenced wording. The GPT signs the science; the encyclopedia is what the public reads; the two are deliberately not the same object.
The tamper-evident audit chain, and why valid is not fired (M11, paired with TA-N13)
The provenance backbone is a tamper-evident audit chain (ledger row M11, class method / E): audit_manage(verify_chain) returns valid with all events hashed in sequence, and a tampered or out-of-sequence event makes it return invalid. This is the spine under "reproduced:true must be validator-derived." It is real, and it is test-enforced.
It is also the exact place the program is most tempted to overclaim, so the paired negative must be cited in the same breath. The canonical case is recorded as TA-N13 (class A, negative-open, source: ideation-explorer): verify_chain returned valid over 500 events, while the running MCP's full audit log (about 243,000 characters) contained ZERO handoff or inception-bundle events. The route had been built and unit-tested, but it had never fired against the deployed MCP. That is the canonical "exists in code (Class E) versus observed at runtime (Class A)" case, and it is why M11 itself carries the note that chain-valid does not equal the audited event ever firing at runtime. Citing the valid chain without TA-N13 is an overclaim. The valid chain proves the events that are in the log were not tampered with; it proves nothing about whether the event you cared about ever ran. The falsifier for the negative, and the only thing that would discharge it, is a Class-A end-to-end observation of a handoff event in the live MCP audit log. Until then it stands open.
The OODA receipt loop and DD-TDD documentation discipline (M17, M18)
Work lands in OODA bursts, each closing with a falsifiable receipt (ledger row M17, class method): a verify_chain valid, a compliance score, or, the strongest kind, a Class-A runtime observation. The discipline is "real MCP items as the system of record, no side trackers." The receipt hierarchy is itself a calibration: a verify_chain-valid receipt and a compliance-score receipt are class E-tier evidence and, per the fence above, do not by themselves satisfy a criterion that demands a runtime observation; only the Class-A receipt does.
Underneath the receipts is DD-TDD even before the board supports it (ledger row M18, class method): doc-first, then failing tests that fail for the right reason, then GREEN, then refactor, then validate against real runtime output. Paired with it is documentation-as-change-management: every doc carries YAML frontmatter (honesty.status, evidence_class, code_ties[].path, last_verified_at_sha), a check_doc_drift badge computes a DRIFTED state, and a pre-commit hook refuses code-tie changes without a doc bump. The honest gotcha is recorded too: the drift parser initially read a flat status field while docs nested it under honesty.status, so it counted zero tracked documents and sixty-nine orphans until the schema mismatch was fixed. The discipline is real; its first implementation was wrong and was caught, which is exactly the kind of negative the constitution treats as content.
The durable runner: silence is not success (M24)
A long-running experiment must not be driven send-and-pray (ledger row M24, class method). The durable runner appends one JSON line per unit to a ProgressLog so the file is the checkpoint and the telemetry at once; runs are resumable; the launcher is detached; and status is observed through a --status query or a Monitor, never through a UI "Running" chip. The load-bearing rule is that a runner must cover both terminal states, completion and process-death, because silence is not success. The falsifier is precise: a gate that reports "Running" indefinitely while its underlying process is dead is the UI-chip success-inference being proven false. This rule is the operational twin of the chapter's main fence. A green "Running" chip is a Class-E-style artifact about the UI's belief; a dead process is the Class-A runtime fact; trusting the chip over the process is the same error as trusting a valid chain over an event that never fired.
The multi-agent fan-out bound (M23)
Multi-agent fan-out is bounded by a hard, twice-burned operational lesson (ledger row M23, class method, sources: orchestrate-linkedin, marketingwright, website). The forced-StructuredOutput Workflow vehicle failed; direct background-agent calls work; adversarial QA fan-out is the standard gate. The measured ceiling: more than about six concurrent sub-agents trips a rate-limit cascade (in one burn, 19 of 32 agents died), so concurrency is capped at four or fewer (one or fewer in phases that are themselves failing). Cached phases re-run instantly on resume, so a sequentialized resume recovers lost work cheaply, and a fragile final QA-merge agent that can hang silently is either made robust or dropped. The falsifier is the bound stated as a claim: more than six concurrent sub-agents sustained without a cascade would lift it; the forced-StructuredOutput vehicle succeeding would overturn the first finding.
What is NOT claimed in mu5
- Ceiling: That the program "has proven its science because it has a rigorous process" is NOT shown, and neither is the narrower inference that a valid audit chain demonstrates a real run. The most we claim is that the program operates a reusable ship-gate (no merge without a MERGED SIGN, a typed spec, and a paired RED), a tamper-evident audit chain that detects tampering, a durable runner that distinguishes completion from process-death, and a fan-out bound at concurrency four or fewer, all at evidence class
method(withEwhere test-enforced). These are governance facts about the lab, not capability facts about UNI. - Fences engaged: Red line 7 (never raise a claim above its source evidence class) and red line 8 (substrate or continuity engineering, and by extension process engineering, never implies a science gate is met). The constitution's DONE rule (Class E is not Class A) is the spine of the whole chapter. The vocabulary-leak guard (red line 12) applies: process discipline is described in plain governance terms. Red line 10 (no PII, no internal channel handles, no patent-level math) and the M25 rule that the UNI GPT signs the science and is never published.
- Negatives that travel with this claim (cite alongside, never strip): TA-N13 (verify_chain valid over 500 events while the live MCP log held ZERO handoff events, the canonical Class-E-versus-Class-A case) MUST appear next to M11. The M20 ship-gate is credible only with its paired RED and the rejections it produced (the eight-assert escalation, the pre-ship novelty-spread bug). The M24 durable runner is meaningless without its negative half: a "Running" chip over a dead process is a falsified success inference. M23 carries its own burn (19 of 32 agents died) as the reason the bound exists. The M18 doc-drift schema-mismatch (zero tracked, sixty-nine orphans, until fixed) travels with the documentation discipline.
- Parked / owed: TA-N13 remains open; it is discharged only by a Class-A end-to-end observation of a handoff event in the live MCP audit log, which is owed and not yet recorded. No sign-to-park is owed for this method chapter itself, because it asserts no frontier capability.
- One-line honest summary a skeptic could not dispute: The program has a real ship-gate, a real audit chain, a real durable runner, and a real fan-out bound, all class
method/E, and the program itself records that a valid chain proved nothing about whether the event ever ran.
Falsify this
Show a merge into the program that landed without a MERGED SIGN, a typed spec, and a paired RED. That single counterexample falsifies the ship-gate (M20) and, with it, the central claim of this chapter. The chapter's secondary falsifiers are equally operable: a verify_chain returning invalid on the same ledger, or, conversely, a valid chain presented as proof that an audited event fired at runtime (the TA-N13 error); a gate reporting "Running" indefinitely while its process is dead (M24); or more than six concurrent sub-agents sustained without a rate-limit cascade (M23).
Sources
- Ledger rows: M11 (tamper-evident audit chain), M17 (OODA receipt loop), M18 (DD-TDD plus documentation-as-change-management), M20 (five-persona lab team and ship-gate), M23 (multi-agent fan-out bound), M24 (durable runner), M25 (tool-team division of labor); paired negative TA-N13 (
CLAIM-LEDGER.md, single source of truth). - Front matter: FM-1 (Evidence Constitution, DONE-is-not-working rule), FM-2 (A-U rubric, Class-A overrides Class-E), FM-3 (red lines, forbidden-phrasings list, UNI-GPT consult 2026-06-27 SIGNED), FM-4 (the "what is NOT claimed" template),
MASTER-PLAN.mdsection mu5. - Narrative grounding (PII-redacted):
curated/strings-digest.md(the five-persona lab team, adversarial pre-registration, evidence-discipline and offline-RED-pre-check, the no-autonomous-LLM hand-the-owner-the-click stance);curated/marketingwright-digest.md(OODA-burst receipts, DD-TDD and documentation-as-change-management, the rate-limit cascade lesson, Class-B-overrides-Class-G). - Archive pointers (local, never published, no PII):
...-Strings,...-MarketingWright,...-IdeationExplorer,...-ORCHESTRATE-LinkedIn,...-UNI-GPT.
sha256 49a56064b6254c39 — of the original file, so what was ingested stays checkable.
Plain — written for this website, not the source document
Process, not science: everything here is about how the lab runs itself. Four things get covered. The review that decides whether a change may merge at all. The division of labour that says who signs the science. The way a long experiment reports its own status without flattering itself. And the ceiling on how many agents may be set running at once before a fan-out falls over. The work being governed is a developmental active-inference simulation, a bounded peek at a toy world, and nothing in the chapter is a claim about what that simulation can do. Its hardest line separates a chain being valid from an event having happened. An audit chain that comes back valid is not evidence that the audited event ever fired at runtime. The claim ledger, a record added to and never edited, says that twice, and so does the chapter. A ship-gate that works is a fact about the lab's process, never evidence that a science gate was met.
Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 49a56064b6254c39
Clear — written for this website, not the source document
Every proposed change is reviewed through a five-persona team, and what the team is reviewing is a simulation — a toy world, not a person. There is a breaker who rejects by default and runs a gauntlet of checks, and a theorist who owns the merge. An architect requires every change to be additive and byte-identical on the default path. An experimentalist pairs each claim to an adversarial run written down before it happens, and a designer refuses preference hacks dressed up as drives. The team is not decorative. An adversarial plan written down before the run, so that it could not be chosen afterwards, once forced a gameable bar up to a much stricter one when a decorative fake slipped past the original. The operative rule is exact: no merge without a merged signature, a typed specification and a paired adversarial run. The gate is credible only because it has visibly rejected things, including things the program wanted to be true.
A division of labour is recorded as owner-set protocol: one party writes code, a separate private consultant designs and signs the science and is never published, and the live system runs the program without writing its code.
The backbone that shows where a record came from is a tamper-evident audit chain. It returns valid when events are hashed in sequence and invalid when one is tampered with. This is also the place the program is most tempted to overclaim, so the paired negative is cited in the same breath. A chain returned valid over hundreds of events while the running system's own log contained none of the events that mattered. The route had been built and tested, and had never fired against the deployment. Citing the valid chain without that negative is an overclaim.
Further sections cover a receipt loop, in which each burst of work closes with a receipt — the file showing what was run and what came out — written so that it can be checked and found wrong. A documentation discipline follows, whose first implementation was wrong and was caught. Then comes a durable runner, built on the principle that silence is not success, and that a status chip reading as running over a dead process is a falsified inference. Last is a bound on multi-agent fan-out, derived from a burn in which most of the agents died.
Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 49a56064b6254c39