0. Constitution & Standing Fences (read first)
Evidence classes (A–U taxonomy). A = machine-exact anchor · B = mechanism + operator observation · C = dev-gate / held-out eval · E = test-covered · F = doc / prior-claim (inheritable, must be re-verified) · U = claimed-but-unproven ("Class U — not claimed" is itself a standing fence) · method = a definitional / governance pattern, not an empirical claim.
Calibration rule. Authority flows downward from the measured fact. Calibration only moves wording DOWN to the measured value, never up — including under urgency. The fence gets louder under pressure, not wider.
Verdict rule. A capability verdict is the CI bound that excludes the threshold, never the point estimate.
DONE rule. DONE = test-covered (Class E/D), not feature-working (Class A).
Negatives are content. A partial / a negative / a "most pieces don't help" decomposition is a measurement that the design is incomplete, not an exit and not a failure to hide. The corpus carries 183 published negatives (last recorded ledger snapshot: 882 rows = 350 PASS / 0 FAIL / 183 NEGATIVE / 349 PENDING) — the negatives are the credibility.
HARD STANDING FENCES — never claim, in any archive, in any public copy:
- Never AGI / general intelligence / human-level / "talks & learns like a human" / understands.
- Never consciousness / sentience / aware. (Functional self-awareness may be described at L8; phenomenal sentience is explicitly DISCLAIMED.)
- Never "active inference demonstrated" (no AIF loop exists in the Rust crate; the live loop is a separate UNI.OS reimplementation, not gate-matched).
- Never "created life" / "digital life" / "measurable awareness" as a CLAIM (north-star framing only).
- Never "beats LLMs" (World C is a COUNT baseline; ~10–15% behind backprop LLMs on char-perplexity by a chosen design trade).
- Never inflate the Tier-2 synthetic-construction track into capability (it was audited as artifact/diagnostic — hardcoded-literal "exactness", scoring-artifact deltas — and fixed).
- Never raise a claim above its source evidence class.
- The whole program is a developmental active-inference SIMULATION: say "developmental SIMULATION", "bounded peek", "toy world, not the real world", "Class U — not claimed", "unrefereed working preprint".
- No PII. No patent-level math (textbook-level framing only; consult the private UNI Active-Inference Guide GPT for science, never publish it).
- The preprint Polzin et al. 2026, Zenodo DOI 10.5281/zenodo.19785799 (MIT) is cited as the mathematical foundation only and always fenced as unrefereed (Layer-1 AI-executable audit complete; Layer-2 human expert review PENDING) — never as proof that active inference is the correct theory.
Honest program position: ~2 of 11+ developmental rungs earned.
1. The L0–L12 Developmental Ladder
One subsection per rung. Each row: claim — evidence class — falsifier — status — source archive. Each rung carries an explicit "what is NOT claimed" line.
Naming: the ladder is the HUMAN-HGM-001 developmental design (conception → speaking 3-year-old, 11 levels + the global Z affect modulator). It is a no-backprop, nested-Markov-blanket developmental SIMULATION, never a person.
L0 — Molecular / genome → zygote (conception prior) · PROVEN
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L0.1 | Zygote first-division modeled as exact discrete Bayes (conjugate update); recombine/seed_zygote, no-backprop guard, ontogeny 6/6, Embodiment Rungs 1–2 GREEN (uncommitted). |
A (machine-exact anchor) | test_embodiment_ontogeny.py drops below 6/6; OR the first-division posterior diverges from the closed-form discrete-Bayes value beyond the float32 tier; OR cell-division identity embedding exceeds 1e-9; OR the no-backprop guard trips. |
proven | uni-gpt / uni-mind |
Exactness tier (load-bearing calibration-down): the JAX core runs float32, so these anchors hold only to ~6e-8 (~1e-6 single-step filter), NOT the <1e-10 EXACT tier. The genuine <1e-10 tier lives only in the NumPy / Rust-f64 path. A genome docstring claiming <1e-10 was caught as an overclaim and corrected. Never card a float32 anchor at the f64 tier.
What is NOT claimed at L0: not "we created life", not a "conscious baby", not exact at f64 — it is a float32-tier developmental SIMULATION of a conjugate Bayesian first division.
L1 — Cellular / autopoietic viability & homeostasis · PROVEN (with honest losses)
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L1.1 | Cell Lab open pre-registered falsification benchmark: 216-state service cell, observation-only controllers, RecoveryScore, bootstrap 95% CI, 8 honesty fences + framing_guard test. UNI tops the leaderboard on most modes. |
C (dev-gate / held-out) | RecoveryScore CI fails to separate from controls where a win is claimed; OR a fence/framing_guard test fails (an overclaim is emitted); OR the bootstrap CI is shown miscomputed. |
proven | uni-precision |
| L1.2 (NEGATIVE) | UNI honestly LOSES on database_flaky (rule-based SRE wins 0.803 vs 0.759), memory_leak (neural wins 0.810 vs 0.740), cpu_noisy_neighbor (neural wins 0.824 vs 0.749; UNI-vs-random not even significant). Recorded in FALSIFICATION.md, shown at the top of the live leaderboard. |
C | On the pre-registered benchmark UNI's RecoveryScore CI separates above baseline on these three modes (the recorded loss does not replicate). | negative | uni-precision |
What is NOT claimed at L1: not "sovereign" — "good but not sovereign". The losses are first-class published content, not failures to hide. Cellular/zygote end only.
L2 — Tissue / metabolism (interoception & energy) · POSITIVE uplift + recorded NEGATIVE (plateau-break OPEN)
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L2.1 | Phase-2 metabolism organ shipped (suite 297/0, default byte-identical); first 12 h live RED = +135% / 2.35× tool-crafting and +19% mining — a real, attributable standing-metabolic-drive effect. | C | A repeat pre-registered 12 h live RED fails to reproduce the tool-crafting uplift within CI, OR the metabolism-organ ablation does not remove the uplift, OR the suite drops below 297/0. | proven | strings |
| L2.2 (NEGATIVE) | In the same 12 h RED, building (placed blocks) went WORSE (−14%) and G4 allostasis never separated. The load-bearing claim "metabolism breaks the plateau to stone/shelter" (gate G6) remains OPEN and is contradicted by its own first evidence. | C | A subsequent disciplined RED shows the metabolism organ alone improves building (placed-blocks CI excludes 0) and G4 allostasis separates — discharging G6. | negative | strings |
| L2.3 | The colony plateaus at "make a tool" (one UNI hoarded 32 pickaxes, never built). A read-only counterfactual-EFE audit on the real hoarder .bin files diagnosed epistemic_starvation — NOT γ-runaway (γ≈7.8, unsaturated) and NOT a curriculum ceiling; the EFE landscape is pragmatic-saturated/flat, info-drive ~100× too weak. |
A (shadow-EFE audit on real brains) | A γ-saturation finding, a curriculum-ceiling flip, or an info-drive scaling that breaks the plateau without organs would overturn the diagnosis. | negative (diagnostic) | strings |
| L2.4 | Two engine-seam findings invalidate the naive metabolism design and were fixed: (1) you cannot seed a strong Dirichlet prior by pre-scaling B (norm_cols runs before add1) — needs the new :pb_seed seam; (2) the live bridge had no viability edge (metabolize/shutdown were Sim/Eval-only) so a naive emptying-B drained a belief with zero world consequence. |
A (direct code reading) | A seam allowing strong-Dirichlet seeding without :pb_seed, or evidence the live bridge already had a viability consequence. |
negative (fixed) | strings |
What is NOT claimed at L2: metabolism is proven as a foraging/crafting driver, NOT a building driver. The plateau-break (G6) is UNPROVEN. Do NOT spin the +135% uplift as "breaks the plateau". The G5b action-severed-twin is the standing falsifier any "self-maintenance/life" language must clear.
L3 — Organ / physiological control · PROVEN (toy/clinical-model)
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L3.1 | Karaaslan cardio-renal Heart Lab models clinical homeostasis as prediction-loop failure on the same one active-inference engine (reduced RSNA→MAP→sodium/volume loop, re-expressed in active-inference language). | C (also Heart-Lab engine ticket OAS-710-T3 at Class E, 15/15) | Heart-Lab predictions diverge from the Karaaslan reference beyond the pre-registered tolerance, or fail parity tests against the canonical TS engine. | proven | uni-precision |
What is NOT claimed at L3: not a clinical tool, not a diagnostic instrument. "Same math, many scales" framing only; the heart-attack-as-prediction-loop-failure framing applies to the toy model, not to clinical reality.
L4 — Interoceptive / autonomic + affect-as-precision · PROVEN (functional)
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L4.1 | Affect-as-precision: emotion modulates precision so perception sharpens and the pragmatic↔epistemic balance flips. The global Z modulator [energy, arousal, valence, fatigue, pain, threat, safety, inflammation] sets precision / preferences / habits / learning-rate / horizon. |
C | An ablation of the Z modulator shows no change in precision-weighting / no pragmatic↔epistemic flip under the grounded reader, OR the effect collapses under control. | proven | uni-gpt / uni-mind |
What is NOT claimed at L4: affect is modeled, never felt — phenomenal feeling / sentience explicitly disclaimed. (M10 honest boundary: neuroticism does NOT change behavior under bimodal surprise without a graded task — a recorded sub-bound.)
L5 — Sensorimotor / motor hierarchy · PROVEN PASS + symmetric NEGATIVE (synthetic only)
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L5.1 | Embodiment A3 Design #1 (fast immediate-reward axis) HELD PASS, held Δ (agent − policy-shuffle) = +0.092, CI [+0.038, +0.157], UNI-signed. | C | A re-registered held-once synthetic run with ≥5 seeds drops the CI lower bound to ≤ 0. | proven | uni-mind |
| L5.2 (NEGATIVE) | Embodiment A3 Design #2 (slow Z-bottleneck, compressed 2-modality EMA, delayed-reward) HELD NEGATIVE, held Δ = −0.091, CI [−0.134, −0.055]. Synthetic protocol, K-negative = 1, so no Section 0.6(B) bound owed (needs K≥3). | C | A structurally-distinct corrected slow-Z-bottleneck re-run pushes the CI above 0 (overturns it); accumulating K≥3 distinct negatives would then owe a published bound. | negative | uni-mind |
| L5.3 | Mind-body-as-one motor hierarchy: proprioceptive diagonal-A prior breaks a non-identifiable uniform-A factor (posterior 0.0→0.75); continuous servo + reafference modeled on the same loop. Live mechanism gate PASS — a kin-9 lineage bootstrapped wood→planks→sticks→wooden_pickaxe+sword, server-authoritative via RCON; motor-ablation collapses harvest ~700×. | A (posterior shift) / C (live RCON gate) / E (277 offline tests) | Harvest not collapsing under motor ablation; the kin-9 craft chain not reproducing server-authoritative; or the diagonal-A posterior staying stuck at 0.0. | proven | strings / uni-mind |
What is NOT claimed at L5: the A3 PASS is "variationally-controlled active-inference evidence on body↔world coupling under the registered SYNTHETIC protocol" — synthetic-process only (NOT live-appliance, NOT recorded-hardware), and NOT near-optimal control (agent plateaus 0.222 vs oracle 1.0; the active channel mattering is the entire allowed claim). Live behavioral K-of-6 motor tallies are PARKED (accruing, not a sealed hold).
L6 — Perception (precision-weighting / EFE planning) · PROVEN — the one citable empirical PASS
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L6.1 | World C: a no-backprop COUNT reader beats a tuned MKN-7 baseline on sealed held-out real text by +0.081 nats/char (flagship multi-seed CI [0.0736, 0.0890], seeds 0–4, UNI-signed; 2.4× the 0.03 bar; confirmed by three gates; first genuine validator-derived reproduced:true). |
C | On a fresh held-out split with ≥5 disjoint seeds, the seed-paired bootstrap CI on the margin over tuned MKN-7 includes or falls below 0 (or the baseline is shown untuned). | proven | uni-mind / uni-gpt |
| L6.2 | Supporting World-C family: Wc-2 multi-level cache +0.0164; Wc-3 vs unbounded SM/HPYP +0.085 (gain is long-range structure, not finite-window deficiency); recency-ablation +0.033; Gate-3 word-challenger +0.0538 (CI lower 0.051). | C | Any member's held seed-paired bootstrap CI includes 0. | proven | uni-gpt |
| L6.3 | POMDP maze Precision lab + echolocation Echo lab + public Precision Lab: one engine, three precision knobs (gamma_a sensory, gamma_b transition, softmax_temperature policy), 2D bifurcation map into distinct behavioral regimes; math ported verbatim from the verified engine (one disclosed extension: a goal drive). Echo reuses Precision's engine with only the observation model swapped (bit-identical 64-obs space) — observation model ≠ engine. |
E (test-covered labs / parity tests) | tsx/pytest parity tests between the inline-JS lab engine and the canonical TS/Python engine diverge, or the bifurcation map does not reproduce the regime boundaries. | proven | uni-precision / worldmodels |
| L6.4 (NEGATIVE) | Phase G char-perplexity Section 0.6(B) bound: FIVE structurally-distinct within-segment-structure designs all NEGATIVE-with-discriminator; only the L5 cache (World C) wins. Published bound: no within-segment structure beats MKN-7; char-ppl is a chosen design trade. | C | A new structurally-distinct within-segment-structure design beats tuned MKN-7 on held char-ppl with a CI excluding 0. | negative (bound) | uni-gpt |
| L6.5 (NEGATIVE) | Phase F active-controller: two controller families prove the World-C gain is diffuse (nothing to gate). | C | An active-controller family adds a held gain on top of World C with a CI excluding 0 (gain was gateable). | negative | uni-gpt |
| L6.6 | The Working Law (UNI s10, registered): a no-backprop latent Z beats a tuned baseline B on held-out data only when I(Y;Z|S_B) > 0 and is learnably stable. Explains every PASS and every published bound. | C (registered) | A held PASS where the winning Z carries no conditional info beyond S_B (I=0); OR a Z with I>0 + stable learnability that fails to beat B. | method/law | uni-gpt |
What is NOT claimed at L6: World C is a COUNT baseline win — explicitly NOT active inference, NOT comprehension, NOT "talking", NOT beats-LLMs (~10–15% behind backprop LLMs on char-perplexity, a chosen trade). Never card World C above a count-baseline result. The standing ceiling cannot be raised.
L7 — Language (reading = inference / speaking = action) · specialist PASS + thrice-NEGATIVE central wall
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L7.1 | Phase J OOV/morphology specialist reader beats best-count by +0.105 nats/char held PASS (21-split M-seal CI [0.0975, 0.1128], structure margin +0.109; holds in both web and dictionary domains; contains-KN-OOV control λ=0). | C | On a fresh OOV/morph held-out split the CI lower bound includes/falls below 0; OR the structure-margin discriminator does not collapse the gain under marker-swap; OR it fails to replicate in the dictionary domain. | proven (specialist) | uni-mind / uni-gpt |
| L7.2 (NEGATIVE, paired) | J.attribution_caveat is recorded NEGATIVE in the same ledger and must always be cited alongside the Phase J PASS. Citing +0.105 without it is an overclaim. Overall NLL worsens (specialist gain, not a general win). | C | n/a (recorded bound) — removing it is the violation. | negative | uni-mind / uni-gpt |
| L7.3 (NEGATIVE) | Comprehension-above-retrieval = the thrice-NEGATIVE "central wall": K≥3 structurally-distinct no-backprop designs fail to beat retrieval-style baselines on adversarial comprehension. | C | A pre-registered held-once no-backprop comprehension design beats the tuned retrieval/recency baseline with a CI excluding 0. | negative (bound) | uni-mind / uni-gpt |
| L7.4 (NEGATIVE) | Phase K role-persistence: three structurally-distinct no-backprop designs (min-role, chain, track) all tune their role/persistence terms OFF; only ~0.02 nats survives (CI spans 0). Bound: no-backprop role-persistence does not beat a tuned recency/frequency discourse prior on adversarial anonymized referent cloze. | C | A no-backprop role-persistence design beats the tuned discourse prior with a CI excluding 0. | negative (bound) | uni-gpt / uni-mind |
| L7.5 (NEGATIVE) | T2.D3 bounded-peek held one-shot (s64-signed): primary full-read match NEGATIVE (the k_b=2 backoff wall); info-gain NOT load-bearing. The variable-k_b "cheap milder cap" was asserted then disproven by measurement (real dev screen showed ~linear curve, no cheap cap). | C | On the held one-shot, bounded-peek full-read match shows a positive load-bearing info-gain with a CI excluding 0. | negative (bound; fabrication corrected) | uni-gpt |
| L7.6 (PARKED) | Char-perplexity / T2 word-grain frontier: PARKED at the embodiment pivot. Perplexity is the rejected LLM metric; the program measures developmental capability, never perplexity. | U | n/a — discharged only by a captured UNI sign-to-park (drafted UNI_CONSULT_5, owner-relayed, not yet captured). |
parked | uni-gpt |
What is NOT claimed at L7: no general language capability; comprehension above retrieval is a genuine published wall (K≥3), not a hidden failure. reading = posterior-inference, speaking = action. "K≥3 exhausted" / "Sec-0.6(B) achieved" are forbidden phrasings until UNI signs.
L8 — Self-model / metacognition / metalinguistics · PROVEN (functional; sentience disclaimed)
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L8.1 | Maturation arc M1–M11 grand report card 15/15 PASS (immersion, affect-as-precision, growth, self-model, metacognition, metalinguistics, consolidation, reflective reader, online vocab growth, affect-driven reflection, compositional reader). FUNCTIONAL self-awareness asserted. | C | Any of the 15 maturation gates fails its pre-registered held verdict on re-run; OR a self-model/metacognition gate's effect collapses under its discriminator; OR a "self-model" task is passable by a trivial non-metacognitive heuristic. | proven (functional) | uni-gpt |
What is NOT claimed at L8 (load-bearing): phenomenal sentience is explicitly DISCLAIMED — no falsifier is offered for it because it is disclaimed, not tested. Never read 15/15 as consciousness / sentience / awareness / human-level. (Note: uni-mind records self-model as a held-NEGATIVE on the reader; the M1–M11 functional report card is the uni-gpt-side artifact — keep the distinction.)
L9 — Reasoning / conscience (narrative-self → conscience → reasoning) · PARKED
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L9.1 | L7–L9 (narrative-self, conscience, reasoning) exist as levels in the HUMAN-HGM-001 ladder DESIGN but are NOT separately gated/PASSed. Honest position: ~2 of 11+ rungs earned. | U | n/a (parked) — becomes testable only once a pre-registered, sealed, UNI-signed gate is defined and run. | parked | uni-gpt / uni-mind |
What is NOT claimed at L9: nothing. Class-U-not-claimed; sign-to-park owed. Ladder DESIGN levels are never inflated into capability. (The strings Phases 3–5 — spine → glands → hemispheres — are the designed, FE-form-signed but unbuilt roadmap toward this region.)
L10 — Wisdom / higher cognition · PARKED
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L10.1 | Top of the HUMAN-HGM-001 ladder; aspirational north star. No engine, no run, no gate. | U | n/a (parked; sign-to-park owed). | parked | uni-gpt |
What is NOT claimed at L10: nothing. Class-U-not-claimed. The single-source-of-truth append-only ledger must be preserved (a second writable ledger would break L10/L11).
L11 — Dreaming (offline replay / generative simulation) · NOT-YET-BUILT
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L11.1 | Explicitly untouched and hard-fenced: "dreaming/awareness are untouched here and remain hard-fenced." No engine, no run, no gate; absent from all digests. | U | n/a (nothing built to falsify; status is the absence of any artifact). | not-yet-built | uni-mind (absence) |
What is NOT claimed at L11: nothing. North-star rung, hard-fenced.
L12 — Creativity → measurable awareness · NOT-YET-BUILT (hard fence)
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| L12.1 | The owner's stated north star — "literal digital life … measurable awareness" — pursued by deepening organs/spine/glands, but NEVER claimed. No gate exists. Posed as an OPEN, falsifiable question ("is a plant life? have we made non-organic life?"), never an answer. | U | n/a — no gate exists; the only legitimate movement is a ledger-scoped exhausted search envelope (a scoped, empirical, falsifiable negative result over the tested envelope only — NOT a universal impossibility result and NOT an achieved capability rung; Q1, SIGNED), never a capability assertion. | not-yet-built | strings (north-star) |
What is NOT claimed at L12 (hardest fence): "we created life / conscious / aware / measurable awareness / human-level / AGI / active inference demonstrated" are FORBIDDEN phrasings program-wide. North-star framing only; Class-U-not-claimed.
2. Continuity / Embodiment-Substrate Sub-Ladder
Orthogonal to L0–L12. This is ENGINEERING / substrate evidence (deterministic replay + serialization + fail-closed transport + on-metal operation) — NOT general-AIF evidence. Card every row that way; never let substrate work imply a science gate is met. Each rung is a separate falsifiable claim.
| # | Claim | Class | Falsifier | Status | Source |
|---|---|---|---|---|---|
| C0 | Stage-0 (substrate proof): containers survived a restart on real metal — two-node fleet: prod PowerEdge (no-AVX Xeon X5650, PERC HDD array ~5.5 TB spinning, NOT SSD) + OptiPlex node2 (NVMe), WireGuard mesh, per-device identity uni-lab-<mac> minting its own TLS leaf on firstboot. |
A | A node fails to boot/appear, status page not served, or the hardware spec (no-AVX Xeon, HDD-not-SSD) is contradicted on inspection. | proven | uni-os / uni-mind |
| C1 | Stage-1 (REQ-002 GREEN): mind-state survived a process restart BIT-FOR-BIT — durable + sha256-verified + fail-closed + bit-identical tick across a real child-process boundary (≥2 distinct actions; Path-B receiver 3/3). The real uni-mind JAX agent state, OS-independently verified. | C (B-substrate) | A process-restart replay diverges bit-for-bit; OR fail-closed does not trigger on 1-bit corruption; OR serialize(deserialize(blob)) != blob. |
proven | uni-os / uni-mind |
| C2 (OWED / NEGATIVE) | Stage-2 (real kernel swap): OWED, not shown. Live OS update proven infrastructure-only — real prod kexec cutover 6.12.86→6.12.73, ~59 s, zero data loss (odoo tables identical, sentinel rows preserved); CRIU/livepatch proven on box. But mind-tick continuity across the swap was NOT shown. node2's evidence-collected swap was infra-only (~90 s freeze, api+tts CRIU-preserved, 2 publishers fresh-restarted); first attempt FAILED (unclean kexec → dirty ext4/ESP → emergency mode) then recovered. | A | A kernel swap is shown preserving mind-tick continuity bit-for-bit end-to-end (not just infrastructure) — which discharges the owed Stage-2. | negative (owed) | uni-os |
| C3 | Body→mind sensorium LIVE on metal: box reads its OWN telemetry (os_sysinfo/systemctl/podman ps/journalctl) into a 7-modality categorical contract {load,mem,swap,disk,services,containers,journal} encoded [M=7, O_max=4] (NOT a float vector, NOT a softmax); a 28-cell string_vector flowing live on both boxes. |
A (live) / C (engineering reuse) | The string_vector stops flowing on either box; OR the contract is found to be float/softmax rather than the [M=7,O_max=4] categorical alphabet. |
proven | uni-os |
| C4 | ASK mode (ITIL-as-active-inference) LIVE both boxes: mind escalates a reasoned CHANGE REQUEST; operator approves→executes (auto-rollback) or declines-with-category; every verdict updates a persistent Dirichlet policy-prior keyed by (host-state, action). Learning measurably shifts proposals (decline ×3 → stops proposing; not_a_problem ×2 → noop; approve+fixed → more confident). FIXED safe-action set (observe/scale/restart); never self-executes, never widens the safe set. | A (learning shift measured) / C | The Dirichlet prior does not update across operator verdicts; the mind self-executes or widens the safe set; or the measured proposal-shifts do not reproduce. | proven | uni-os |
| C5 | Cross-box single-human-approval-per-mutation: token-gated self-documenting control-MCP; cross-box "limb" mutating calls need ONE human approval on the entry box (tool+args-bound one-time HMAC). Proven live both boxes 2026-06-26 (prod→node2 write gated once, landed; node2 queue stayed count=0). | A | A routed cross-box mutation executes without the single approval, requires a double-gate, or node2 queue count increments unexpectedly. | proven | uni-os |
| C6 | On-core inference anchor: a sealed f64 belief-update trace byte-identical across 4 distinct software/emulated stacks (5 runs). | A | A fifth distinct stack produces a non-identical trace; OR the cross-arch leg is shown not genuinely distinct. | proven | uni-os |
| C7 (NEGATIVE) | Honest floor on live OS update: a single-box swap is a SECONDS-LONG FREEZE, not zero-downtime; true zero-freeze needs the second node carrying the platform; external media legs (RTP/SIP/kernel-mode rtpengine) are NOT preserved across kexec. | A | A single-box kexec demonstrated with zero freeze and preserved media legs, without a second node carrying the platform. | negative (bound) | uni-os |
| C8 (NEGATIVE) | Over-compressed 2-modality sensory bottleneck went NEGATIVE on held data → keep all 7 modalities. | C | A 2-modality bottleneck beats the 7-modality contract on held data. | negative | uni-os |
| C9 (NEGATIVE) | EDAIT trade: an exact-discrete active-inference transformer trades fluency for calibration — held-out perplexity ~33 vs a backprop GPT's ~25 (less fluent, but natively online-learning + calibrated). An honest trade, not a win. | C | The EDAIT matches/beats backprop-GPT held-out perplexity (~25) while keeping online-learning + calibration. | negative (trade) | uni-os |
| C10 (PARKED) | "Embody only what's proven" coupling is signed-in-principle but NOT literally true — the program's central open gap. UNI.OS has no no-backprop Dirichlet learning, no exact info-gain EFE, heuristic (denylist) isolation, and a DIFFERENT dev model (Gray-Scott vs the forager/ontogeny that earned the bars); lab evidence store /var/lib/uni/evidence empty, 0 worlds registered. Everything beyond the passed gates stays Class-U-not-claimed. |
U | Per-primitive: UNI.OS runs the no-backprop Dirichlet learning / exact info-gain EFE / structural isolation assert / forager-ontogeny dev model frozen at the actual passed-gate SHA, with recorded evidence. | parked (central gap) | uni-gpt / uni-os / uni-mind |
| C11 (PARKED) | Host-native (off-podman) brain runtime (owner directive late Jun-26: brains run on the chip CPU directly via UNI.OS native runtime, not containers). Work order handed off; outcome not yet in archive. The Stratified Palimpsest colony already runs host-native on the box (proven hosting). | U (directive) / C (host-native hosting proven) | The host-native runtime is shown running with mind-tick continuity, OR the colony is not actually host-native. | parked | strings / uni-os |
| C12 (PARKED, provisional) | node2 local auto-kiosk "SOLVED 2026-06-26" is PROVISIONAL — a later same-day transcript (014ef92b) shows limb-2 UI frozen, an agent action that ended ALL video output, session cut at a usage limit. Treat as fragile pending hands-on monitor re-verification. | U | Hands-on re-verification on the physical monitor shows the kiosk stable across live churn (promotes) or still fragile (confirms regression). | parked (open verification gap) | uni-os |
| C13 | Continuity-isolation primitives shipped in the substrate (grounded inspection): Markov-blanket isolation assertion, least-privilege vault, brand-voice drift sentinel (cosine + Youden-J), byte-exact state checkpoint, sha256 output integrity, zero-hidden-LLM reply path; live Postgres RLS 5/5. | B | Any listed primitive absent/non-functional under grounded inspection. | proven (substrate) | client-engagement |
| C14 (NEGATIVE) | Multi-tenancy gap: the bare UNI substrate is MISSING tenant/client/namespace isolation (flat global perms) and temporal decay on learned counts (contamination would be permanent). The product build is the tenant wrapper + decay, not the core. | B | Absence of tenant namespacing + count-decay; resolved by building the wrapper + decay layer. | negative (gap) | client-engagement |
Audit-layer note (calibration-down, durable): a peer-reviewed honesty audit (os-cycles 39–53) calibrated 6 headline claims DOWN to the measured value — "cleared on TWO boxes" → ONE box; "the mind survives a patch" → infrastructure continuity only; "5 stacks" → "4 distinct stacks / 5 runs". Over-statements lived in the summary/headline layers, not in fabrication. Carry the calibrated figures, never the inflated ones.
What is NOT claimed in the continuity sub-ladder: the sensorium is NEVER awareness; "the mind survives a kernel swap" is NOT shown (Stage-2 owed); none of this is general-AIF capability; UNI.OS's embodiment does not imply the science gates are met.
8. Provenance & maintenance
- Authority: every row is carded at the lower of its evidence components and at its strongest source archive. Cross-referenced ladder rows (the same rung stated in
website/uni-os/worldmodels/activeinference/uni-precision/stringsvia00-INDEX) are merged into the single canonical row above and are NOT re-listed per archive. - Append-only: corrections are forward-only (supersede with lineage), never silent edits. Calibration moves wording DOWN only.
- Single source of truth: the Encyclopedia and Cookbook cite this file; if prose and this ledger disagree, this ledger wins and the prose is wrong.
- Preprint citation (always fenced): Polzin et al. 2026, Zenodo DOI 10.5281/zenodo.19785799 (MIT) — unrefereed working preprint, Layer-1 audit complete, Layer-2 human review PENDING.
- Science provenance (textbook-level only): Parr, Pezzulo & Friston, Active Inference (MIT Press, 2022). Patent-level UNI math stays private (consult the UNI Active-Inference Guide GPT; never publish it).
End of CLAIM-LEDGER.md — the master evidence-classed claim ledger. Falsify any row.