UNI Universal Natural Intelligence

Wiki · The Cookbook

UNI-GPT Consult — 2026-06-27

The Cookbook · cookbook/UNI-GPT-CONSULT-2026-06-27.md @ 575fc93d9d31 (main) — opens the published snapshot e850f872196d

How to read this page

Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.

The Cookbook is the method carried out step by step: 34 pages of recipes for building a developmental active-inference SIMULATION — a bounded peek at a toy world, never a person. The front matter says that word is never softened under any pressure, so it is not softened here. The recipes run from the molecular and cellular rungs up through metabolism, motor control, perception, language and metacognition, and on to rungs that are still open questions. Around them sit a set of kitchen rules, a shared pantry of engines and primitives, and a second family of recipes about nature itself — rocks, water, air, stars, DNA, ants, whales, bats, humans.

It is for the reader asking what building this would actually take. Each recipe names its ingredients, the order of work, the tests to run at that stage, and the point at which a step stops being something already carried out and becomes something proposed.

Begin with the front matter and then the kitchen rules. Those two pages fix the honest position and the fence labels that every later recipe leans on, and without them the status markers on a recipe are easy to skim past. After that the recipes can be read in any order.

The nature recipes sit slightly apart and should be read that way. They cite outside science — geology, chemistry, biology, astrophysics — and a nature citation is never a UNI gate: those chapters contain zero UNI claims and raise no rung.

What it is not: a claim that the whole ladder has been cooked. The book recommends the complete recipe and, on the same page, labels every rung by its real state — that tension is deliberate and is the thing the book is built around. Where a recipe and the claim ledger disagree, the ledger wins.

Your browser cannot switch reading levels, so the document itself is shown.

Precise — the source document

This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.

Consultant: UNI Active Inference Guide custom GPT (GPT-5.5 Thinking), by Michael Polzin. Driver: Claude Code via Claude-in-Chrome bridge. Purpose: confirm the Session-2 master plans (encyclopedia + cookbook); answers fold back into the plans calibrated DOWN only (never up). Verbatim capture; UI chrome trimmed. The ledger remains the single source of truth — a GPT answer that would raise a claim is recorded but NOT applied.

Context block sent first (honest position: developmental SIMULATION, ~2 of 11+ rungs; fences held; L0-L8 proven w/ NEGATIVEs at L2/L5/L7, L9-L10 parked, L11/L12 not-yet-built). GPT acknowledged: "I can read you; UNI Active Inference Guide is ready for Q1–Q7."


Q1 — Will UNI sign the parked frontier as a published exhausted bound, and with what exact wording?

Verdict: SIGN-WITH-CHANGES.

Signs the park, but will NOT sign the phrase "published exhausted bound" unless immediately qualified as a ledger-scoped, implementation-scoped, data-split-scoped negative result / exhausted search envelope — not a universal theorem and not an achieved capability rung. ("Exhausted bound" can sound like mathematical exhaustion of a design space; what we actually have is a published negative frontier result — the tested no-backprop route did not clear the chosen frontier under the registered protocol. Stronger and more honest, but it must stay empirical and scoped.)

Canonical signed wording (L7/T2/L9–L10 Park Statement):

We park the L7 char-perplexity / T2 word-grain frontier and the L9–L10 role-persistence ladder as a ledger-scoped exhausted search envelope, not a universal impossibility result. Under the recorded corpus, splits, metrics, implementation, compute budget, ablation set, and comparison baselines in the ledger, no tested within-segment no-backprop structure improved char-perplexity beyond the tuned MKN-7 count baseline. This establishes a negative bound over the tested envelope only. It does not establish that all K≥3 structures are exhausted, that no future within-segment model can improve, or that any broader language-modeling frontier has been closed.

Char-perplexity is recorded as a chosen design trade, not a failed claim of general language superiority: World C remains a count-model baseline — not an active-inference demonstration, not a beats-LLMs claim, not evidence of human-level language capacity.

For L9–L10: the tested no-backprop role-persistence mechanisms did not beat a tuned recency-frequency discourse prior under the ledgered protocol. This parks the rung as an honest negative frontier result. It does not achieve Sec-0.6(B), does not demonstrate active inference, and does not license claims of metacognition, consciousness, AGI, human-level discourse, or created life.

The recipe may describe the attempted mechanisms and the negative result, but the ledger is the single source of truth. If recipe and ledger conflict, the ledger wins.

Phrases explicitly banned → replacements:

  • "K≥3 exhausted" → "The registered tested K conditions did not reverse the result."
  • "Sec-0.6(B) achieved" → "Sec-0.6(B) remains unearned / parked pending a future result that beats the registered discourse prior under ledgered evaluation."
  • "UNI demonstrates language/metacognition at L9–L10" → "UNI records a negative L9–L10 frontier test under the no-backprop developmental simulation program."
  • "This proves no no-backprop model can beat MKN-7" → "No tested no-backprop variant in the registered envelope beat MKN-7 on the chosen char-ppl metric."

8 conditions for the park to remain honest (signed only if all hold):

  1. Ledger completeness — name exact corpus, train/val/test splits, preprocessing, tokenization/char inventory, smoothing, MKN-7 tuning method, candidate variants, K values tested, seeds, stopping criteria.
  2. Baseline fairness — MKN-7 and the recency-frequency prior tuned within the same declared budget as the candidates, or disclose the asymmetry.
  3. Metric containment — char-perplexity is the chosen metric/design trade, not a total language-capability measure; cannot imply understanding/role-persistence/ metacognition/active inference.
  4. No universal-quantifier leak — never drift from "no tested variant under this protocol" to "no possible variant" / "K≥3 exhausted" / "frontier theoretically closed."
  5. No rung inflation — L9–L10 stay parked; L11 dreaming + L12 not-yet-built and hard-fenced; no "earned rung" language without a ledgered criterion + passing evidence.
  6. Evidence-class labels stay attached — negative result is a ledgered empirical result for the tested envelope; any extrapolation is Class C at best.
  7. Recipe subordinate to ledger — the Cookbook may teach route + failure mode but must state the ledger supersedes prose.
  8. No active-inference-demonstration claim — World C is a count baseline and the parked L9–L10 tests did not clear the discourse-prior comparison, so the park cannot be described as "active inference demonstrated."

Q2 — Which missing UNI.OS primitive to port first (gap C10)?

Verdict: port no-backprop DIRICHLET LEARNING first, frozen at the actual passed-gate science-repo SHA. Vendor/import the exact primitive (do NOT re-author, reinterpret, or "improve" it during the port); record the SHA; prove byte/behavior equivalence to the frozen science implementation.

Ranking of the four (highest→lowest first-port value):

  1. Dirichlet learning — highest. Central, canonical, no-backprop, low-risk, directly testable. Converts UNI.OS from "runs fixed AIF-shaped control" to "embodies a core no-backprop learning primitive." Makes the fast-inference / slow-learning separation physically present, not merely described.
  2. Exact info-gain EFE — highest conceptual importance, but second: if ported before learning it can still leave UNI.OS a fixed-model planner; doesn't discharge the "no-backprop developmental substrate" gap as strongly.
  3. Structural whitelist isolation — highest safety/architecture value, not highest science-port value; it's a guardrail primitive (proves boundary discipline, not learning/epistemics). Implement soon, but not first.
  4. Dev-model alignment (replace Gray-Scott with forager/ontogeny) — necessary before claiming substrate equivalence, but too high-cost/high-confound to do first; port the learning primitive first for a clean falsifiable bridge.

The math interface (must be reproduced exactly):

  • A-learning (conjugate count): a_ij <- a_ij + sum_tau o_tau,i * s_bar_tau,j
  • Expected-log used in place of raw ln A: E_Q[ln A_ij] = psi(a_ij) - psi(sum_k a_kj)
  • Preferred: B/D/E update only by the same registered conjugate-count family (transitions, initial-state prior, habit/policy prior) — never optimizer steps.

Discharge condition (frozen-SHA equivalence gate, in CI + ledger): given the same initial Dirichlet concentrations, observations o_tau, beliefs s_bar_tau, action trace, and learning-rate as the science passed-gate fixture, UNI.OS produces the same concentration updates and same expected-log tensors within declared tolerance, with no gradient descent / backprop / learned weights outside the Dirichlet update. Fixtures: A-learning, expected-log, optional B/D/E, a no-backprop guard (test FAILS if any autograd/optimizer path runs or a trainable weight changes outside the concentration tensors), and a frozen-provenance record (primitive name, science SHA, passed-gate test id, fixture hash, port commit, tolerance, deviations).

Falsifier: not discharged if UNI.OS updates A/B/D/E by any heuristic/decay/ optimizer/learned-embedding/hidden-gradient/hand-normalized-frequency path that doesn't reproduce the frozen primitive's outputs — OR if counts update correctly but planning/perception still read raw normalized A where the gate requires E_Q[ln A] (looks like learning while failing the real interface — especially dangerous).

What stays Class-U-not-claimed until the other three land:

  • NOT "UNI.OS implements full EFE-based active-inference planning" (heuristic planner gap remains until the passed-gate EFE primitive is ported).
  • NOT "UNI.OS structurally enforces the Markov blanket" (isolation is denylist/ heuristic until the whitelist assert lands).
  • NOT "UNI.OS embodies the passed developmental arc" (still Gray-Scott, not the forager/ontogeny that earned the bars).
  • NOT "active inference demonstrated" — even after Dirichlet learning lands, the honest status is "one passed-gate no-backprop learning primitive is literally embodied," nothing broader.

Recommended ledger wording: "C10 partial discharge — Dirichlet learning port. UNI.OS now literally embodies the frozen passed-gate no-backprop Dirichlet learning primitive from science repo SHA , verified by equivalence tests over concentration updates and expected-log tensors. This discharges only the learning-primitive gap. Exact info-gain EFE, structural-whitelist blanket enforcement, and forager/ontogeny developmental-model equivalence remain unported and unclaimed."


Q3 — Smallest cure to break the L2 building plateau (gate G6)

Verdict: SIGN-WITH-CONDITIONS. The smallest structurally-distinct cure is a build-affordance epistemic micro-organ — NOT a gamma change, NOT a second metabolism organ, NOT a placed-block reward bonus.

The change: add one building-specific hidden factor + one policy term valuing information gain about where a block can usefully be placed next.

  • organ: build_epistemic_frontier
  • hidden factor: z_build in {unknown_placeable, placeable_support, shelter_contributing, blocked/useless}
  • drive: maximize expected info gain over z_build for candidate inspect/move/place policies
  • scope: active only when shelter/stone plateau preconditions are near but not achieved
  • Added epistemic term per policy: G_new = G_old - beta_build_IG * E_Q[ D_KL( Q(z_build|o,pi) || Q(z_build|pi) ) ] (or ambiguity/risk form G_new = G_old + lambda_build * H[A_build] s_pi, sign consistent with impl).
  • Scale β by matching, not hand-waving: beta_build_IG = median(|Δ pragmatic G for food/tool|) / median(|Δ build IG term|); clamp initial to {25x, 100x}; 100x only if offline RED shows 25x underpowered.

(b) Why epistemic, not gamma: the diagnosis was info-drive ~100x too weak with gamma ~7.8 unsaturated. Gamma is a precision over G — Q(pi)=sigma(ln E - gamma*G_pi - F_pi) — so raising gamma only sharpens the existing ranking; if build-relevant info gain is missing/100x too small, gamma sharpens the wrong ranking. The cure is to make build-relevant uncertainty part of what the planner can value, not to make the planner more decisive. (Organ metaphor stays Class-C design language; the standard AIF part is only the EFE decomposition + policy selection.)

(c) Paired RED — G6_BUILD_EPI_FRONTIER_PAIRED_RED_v1:

  • Matched kin-pair seeds. Control = kin-N+1 (current organ/gamma/build machinery, no frontier term). Treatment = kin-N (identical + build_epistemic_frontier, β fixed from offline RED, gamma unchanged). No other cures.
  • Offline RED pre-check (replay prior traces, read-only) — pass only if all hold: build-policy rank shift appears (P(place/inspect-build in top-k) treatment>control); gamma non-diagnostic (unchanged, unsaturated, effect not dependent on raising gamma); the info term is causally responsible (ΔG_build_IG explains the shift, not Δrisk_food/ Δtool/tie-breaking); foraging/crafting not cannibalized (within non-inferiority). Suggested bar: ≥25% relative increase in build-relevant policies entering top-3, food/tool rank within ±5%, gamma unsaturated.
  • Live paired RED: unit = matched kin pair; same 12h, same world dist, paired seeds, no peeking-based tuning. Primary: placed_blocks_delta = T - C, 95% paired bootstrap CI excludes 0 (positive). Co-primary (G4): allostasis_index delta CI excludes 0 in intended direction, no viability collapse. Use the existing ledgered G4 metric — do NOT invent one after the run. Secondary: stone/shelter progress >0, crafting/mining non-inferior, distance-to-shelter improves.
  • Required discriminator/ablation (gain MUST collapse): IG-zero (beta_build_IG=0), shuffled-affordance (z_build permuted across sites), gamma-only (no organ, gamma matched). Credited only if the gain appears with build IG present and collapses when the IG channel is zeroed/scrambled, while gamma-only fails to reproduce it.

(d) Falsifiers: placed-blocks CI includes 0 / negative; OR blocks improve but G4 never separates (construction activity, not plateau-break allostasis); OR the ablation fails to collapse the gain (not caused by build epistemics); OR gamma-only reproduces it (the "not gamma" diagnosis was wrong); OR blocks bought by damaging survival/tooling beyond non-inferiority (trade-off/pathology, not a clean G6 discharge). If offline RED predicts no rank shift but live improves → log "behavioral improvement observed; mechanism not proven."

Status: until the RED clears both bars and the ablations collapse the gain, this is a Class-C design hypothesis supported by a read-only counterfactual diagnosis, NOT an achieved plateau-break.


Q4 — The first sealed, falsifiable L9 gate

Verdict: SIGN — gate name L9-G1: Cavity-correct delayed commitment-state prediction. (Caveat: the GPT noted "Phase-3 spine"/"cavity principle" aren't in its source docs, so it treats them as our design labels and grounds the answer in the standard hierarchical/deep active-inference pattern — slow high levels set empirical priors for fast low levels; a modeling/inference claim, NOT a claim about reasoning or conscience in the world.)

(a) Smallest L9 target (NOT reasoning/conscience/narrative-self): build the Phase-3 spine to carry ONE slow latent z_spine = {role_id, active_commitment, norm/constraint_tag} across several lower-level segments, then test whether it improves prediction when the local segment points the wrong way. Sealed episode: seg-1 establishes a commitment; segs 2..N are distractors/locally-plausible alternatives; final segment requires predicting the next action / conflict cue.

  • Primary bar: ΔNLL = NLL_baseline − NLL_spine > 0, 95% paired bootstrap CI excludes 0; effect-size guard ≥3% relative NLL reduction on the load-bearing delayed-commitment subset.
  • Secondary: conflict/violation tag AUC or F1 +≥0.05 absolute, paired CI excludes 0.
  • Calibration guard: ECE worsens by ≤ a pre-registered margin (e.g. +0.02).
  • Smallest because it needs only persistent cross-level state + constraint-sensitive prediction — no open-ended reasoning, morality, selfhood, language, or metacognition.

(b) Tuned baseline it must beat: a tuned recency-frequency discourse prior (segment-local predictor + exp-decayed role/action frequencies + low-order transition counts over role/action/conflict tags + tuned window k, decay τ, smoothing α) — may use recent tags/context but may NOT carry a slow latent spine with cavity-correct cross-level residuals. Tuned on train/val only, frozen before test, reported with its own ablations. Beating only an untuned recency model leaves the gate OPEN. (Deliberately the same baseline as the parked L9/L10 NEGATIVE — the spine earns only the narrower "beats the tuned recency-frequency prior on delayed-commitment prediction.")

(c) Discriminator + computed-residual ablation:

  • Load-bearing discriminator (local-distractor commitment reversal): early seg sets commitment C; middle segs make a different action A_local locally frequent; final seg requires predicting commitment-consistent A_commit or conflict-if-A_local. Credit only if the improvement concentrates on THIS split, not easy same-topic continuity.
  • Computed-residual (cavity) ablation — not "turn off the magic string": r_t(x) = log p(x_t | q_high(z_t)) − log p(x_t | q_cavity(z_t)) where q_cavity excludes the receiving segment's own evidence path (so the same evidence isn't counted twice). Treatment uses q_local + r_t; ablation sets r_t := 0 (or episode-shuffles r_t after computation). The ablation MUST collapse the gain on the discriminator split.
  • Double-counting sentinel — naive-spine ablation: use the full high-level posterior as a downward prior WITHOUT cavity subtraction; expected to sharpen confidence but worsen calibration / fail the no-double-counting residual test.

(d) Fenced pass-claim wording (signed): "L9-G1 passed, fenced. In a sealed delayed-commitment prediction task, the Phase-3 spine improved commitment-conditioned next-action and conflict-tag prediction over a tuned recency-frequency discourse prior; the improvement concentrated on the pre-registered local-distractor split and collapsed when the computed cavity residual was zeroed or shuffled. Narrow claim: the system can carry a slow cross-level prior as a computed residual without double-counting and use it to improve a specific role/commitment-persistence prediction task." Non-license: does NOT claim reasoning, conscience, narrative-self, metacognition, awareness, consciousness, AGI, human-level language, created life, or "active inference demonstrated"; does NOT unpark L9 globally. Evidence class: hierarchical AIF substrate Class B/Sec; Phase-3 spine design Class C; passed gate = ledgered empirical result for this simulation only.

(e) Falsifiers (gate fails if any): ΔNLL CI includes 0 or reduction <3%; improvement only on easy continuity (not the distractor split); zero/shuffle of r_t does NOT collapse the gain; a naive non-cavity spine matches/beats the cavity spine without a calibration penalty; a stronger tuned recency/transition baseline closes the gap; leakage (labels/ templates/episode-IDs/handcrafted norm strings encode the answer); the gain trades off against already-earned lower rungs; OR any write-up implies reasoning/conscience/ metacognition/AIF-demonstrated/AGI/consciousness/human-level/created-life.


Q5 — What an L11 "dreaming" gate would measure (never "dreamed")

Verdict: SIGN — gate L11-R1: held sparse-sequence retention after offline replay. Offline replay is a UNI design primitive / Class-C experimental mechanism, not a new established AIF result; the grounded machinery is ordinary POMDP A,B,C,D,E with a strict model/process split.

(a) Offline-replay mechanism (no psychological language): a sealed consolidation interval where the external observation channel is detached (Z_offline=1) and the agent's own generative model is run forward to produce temporally coherent replay/sample trajectories. Consolidation writes ONLY through pre-existing ledgered update rules — no new optimizer, no backprop, no hidden supervised labels, no live observations. Online phase Z_offline=0 (real o_t in). Offline phase Z_offline=1: block observation injection; sample start states from stored posteriors / registered replay seeds; sample policies from frozen Q(pi)/E or a pre-registered replay policy; roll states through B; sample synthetic outcomes through A; consolidate via existing rules. Durable runner checkpoints {replay seed, model hash, Z, replay count, update deltas}; resume must reproduce the same trace + deltas. Observation-detach is load-bearing — generated ô is model content, not new external data.

(b) Smallest gate (predictive before behavioral): paired kin seeds; both arms get identical online exposure containing sparse low-frequency multi-step contingencies (cue/context → action_1 → hidden transition → outcome). Offline interval: treatment Z_offline=1 coherent replay enabled; control = same checkpoint / same runner path / same no-new-data, but no replay trajectories and no replay-derived updates. Held eval: no learning, probes are NOT replayed traces. Primary: held_sequence_NLL = -log Q_agent(correct held next outcome/action | probe); pass iff paired bootstrap 95% CI for ΔNLL = NLL_control − NLL_treatment excludes 0 (positive) AND point estimate ≥ δ = max(0.02 nats/decision, 2% relative). Secondary (if action exists): first_action_success_delta>0, CI excludes 0. Primary stays held predictive NLL — the smallest downstream effect showing consolidation changed the usable model.

(c) Discriminator + collapse ablations: discriminator = held probes where marginal frequency MISLEADS — online sees common A→D plus rare task-relevant A→B→C→reward; held probe starts at A in the context where A→B→C is correct while recency favors A→D (control overweights recency; treatment improves the coherent sequence). Gain MUST collapse under: (1) shuffled temporal replay (same counts, transition order broken); (2) random replay (same compute, trajectories from marginals not coherent B rollouts); (3) no-write replay (coherent trajectories generated, consolidation writes disabled). Safety ablation: observation-attached offline interval — if THAT wins, the result is extra online data, not replay. Credited only if coherent model-generated replay improves held sparse-sequence prediction AND collapses under shuffle/random/no-write.

(d) Fenced pass-claim (signed): "L11-R1 passed, fenced. In a sealed paired evaluation, an offline generative-replay consolidation interval improved held sparse-sequence predictive NLL over a no-replay control; used no live observations during replay; and collapsed under shuffled/random/no-write replay ablations. Narrow claim: the simulation has a measurable offline-replay consolidation effect under the registered protocol." Non-license: does NOT claim consciousness, awareness, sleep, human-like mentation, AGI, human-level cognition, created life, reasoning, creativity, or "active inference demonstrated"; does NOT clear L11 globally. Class: POMDP substrate B; offline-replay design C; passed gate = ledgered empirical result for this simulation only; psychological/biological interpretation NOT claimed.

(e) Falsifiers (fails if any): ΔNLL CI includes 0 / negative; CI excludes 0 but point estimate < δ; shuffled/random replay does NOT collapse the gain (coherence not load-bearing); no-write replay still improves (benefit isn't consolidation); offline channel not actually detached (any live obs / env query / supervised label / probe leakage in the interval); unfair control (more data / different wall-clock / extra tuning / different checkpoint); gain is probe leakage (replay contains exact probes); lower-rung regression beyond non-inferiority; durable runner non-reproducible on resume; OR claim inflation implying consciousness/sleep/human-mentation/AGI/created-life/global L11.


Q6 — Next structurally-distinct L5 motor design (toward live PASS or K>=3 bound)

Verdict: SIGN — accrue Design #3: Proprioceptive Servo Bridge, aimed at a LIVE PASS (not bound-seeking; don't design a strawman to farm negatives). Class-C UNI design; the standard part is the AIF control vocabulary (sensors as likelihood channels, proprioception as a modality, precision as gain, hybrid discrete-over- continuous control).

(a) Design #3 — changes 4 dimensions vs #1/#2:

  1. Coupling topology: scalar reward/Z coupling → closed-loop triadic (high-level task policy → motor setpoint → proprioceptive error → corrective action).
  2. Timescale source: immediate-reward (#1) / slow Z-bottleneck (#2) → event-triggered fast motor correction (update when proprioceptive PE crosses threshold, else hold the current motor primitive).
  3. Information bottleneck: global Z → low-dim proprioceptive residual r_motor = desired_pose/velocity − inferred_pose/velocity.
  4. Control path: direct discrete action bias → a motor shim converting policy intent into continuous/local corrective commands. Pipeline: high-level action → motor setpoint → ε_prop → servo bridge → world action.

(b) Aim = live PASS: #1 already showed a synthetic held pass on the fast immediate-reward axis; #2 a symmetric synthetic negative on the slow Z-bottleneck. Highest-EFE next move is to ask whether the positive motor signal SURVIVES a more embodied, live-relevant closed-loop mechanism — not to farm another negative. If it fails cleanly it may count as K-negative=2 toward a future motor bound, but only if truly structurally distinct and all pre-registered bars/controls valid.

(c) Protocol L5_D3_PROPRIO_SERVO_BRIDGE_HELD_v1: paired kin seeds; treatment = servo bridge enabled; control = best current L5 stack without it (identical high-level policy / reward / env / compute / eval window). One cure at a time (no Z-bottleneck cure, no new learning rules, no exploration bonus, no reward edits).

  • Held bar: held_motor_delta = treatment − control, paired bootstrap 95% CI excludes 0 positively, point estimate ≥ +0.05 (smaller than #1's +0.092 but nontrivial).
  • Live bar: paired CI for live motor-success delta excludes 0 positively AND no earned survival/allostasis/task metric regresses beyond non-inferiority.
  • motor_score = task-relevant successful actions − collision/overshoot/oscillation − stalled-control − energy/recovery penalties. ("More movement" must NOT count as success.)
  • Discriminator (delayed/slipped embodiment trials): actuator noise / one-step action delay / friction change / positional perturbation / contact-state ambiguity. Signature: a #1-like fast reward axis improves direct choices but degrades under slippage/delay; #3 recovers because the proprioceptive residual drives correction.
  • Ablations that must collapse the gain: (A) ε_prop := 0 (shim present, compute runs); (B) ε_prop shuffled across time/kin (mistimed correction); (C) open-loop setpoint only (mapping kept, no feedback correction). Credit only if the gain appears on perturbation trials and collapses under residual zero/shuffle.

(d) Fenced claim wording:

  • PASS: "L5 Design #3 sealed PASS, fenced. In the registered held/live paired protocol the Proprioceptive Servo Bridge improved motor-control performance over the tuned control stack (paired CI excludes 0); the gain concentrated on closed-loop perturbation trials and collapsed when the computed proprioceptive residual was zeroed or shuffled." Non-license: NOT general motor intelligence / human-like embodiment / AGI / consciousness / created life / "active inference demonstrated"; does NOT erase the Design #2 negative.
  • NEGATIVE-toward-bound: "L5 Design #3 sealed NEGATIVE, fenced. ... failed to improve the motor metric (CI did not exclude 0 positively). Because it changed coupling/timescale/bottleneck/control-path vs #1/#2, it may count as one structurally distinct negative toward a future K≥3 motor bound, provided all validity checks passed." Non-license: NOT a motor impossibility result; with #2+#3 the ledger has at most K-negative=2; no Sec-0.6(B) motor bound owed until K≥3.

(e) Falsifier / non-informative-K: PASS falsified if CI includes 0 / negative; gains only on easy non-perturbed trials; residual zero/shuffle does NOT collapse the gain; open-loop-only matches treatment; OR motor score improved by damaging survival/allostasis/energy/task. Does NOT count toward K≥3 if: <2 structural dimensions changed; undertuned/less compute vs control; silently combines cures; protocol so different that deltas aren't comparable to #1/#2; discriminator absent/too easy; ablation is hardcoded (disables whole motor stack) not causal (only ε_prop); bug/seed-imbalance/env-drift/logging-failure breaks paired inference; or result depends on post-hoc metric selection.


Q7 — L12 published-bound language + the overall sign-off

(a) Exact public-facing L12 bound/framing (signed):

Open L12 question, not a claim. UNI's L12 frontier asks a falsifiable question: Can a non-organic, no-backprop developmental active-inference simulation ever produce measurable awareness-PROXY behavior that both matches or beats the best contemporary LLM baselines on the same sealed tasks AND shows substrate-distinct properties those LLMs do not show? Today the answer is not known, not claimed, and not implied. UNI has not created consciousness, human-level intelligence, AGI, or life. "Awareness" here means only a future pre-registered proxy suite — calibrated uncertainty, global availability of information across modules, self-monitoring that improves later policy, contradiction detection between report and trace, counterfactual access, and calibrated abstention. Passing such a suite would license only: "UNI passed specified awareness-proxy tests under the registered protocol." It would NOT license: "UNI is aware," "UNI is conscious," "UNI is alive," or "we created non-organic life."

Plant-life clarification: plants are biological life in the ordinary sense; UNI does not use that to claim a simulation is alive. The published UNI question is narrower: whether a non-organic developmental simulation can meet pre-registered life-like persistence and awareness-proxy tests without collapsing into metaphor, declaration, or LLM comparison games.

(b) Forbidden phrasings (ban outright): "UNI is conscious / aware / self-aware / has measurable awareness / has a mind / has a conscience / reasons like a human / is human-level / is AGI / created life / created non-organic life / is alive / is a synthetic organism / dreamed / is creative in the human sense / demonstrates active inference / proves the FEP creates minds / proves plants and UNI are the same kind of life / beat consciousness tests / beat LLMs therefore awareness / passed L12 / L12 achieved." Replace with: "UNI passed a specified proxy test / improved a registered downstream metric / matched-or-exceeded the registered LLM baseline on this sealed task / showed a substrate-distinct effect under this ablation / remains a developmental active-inference simulation / the awareness question remains open."

(c) Proposal-entry checklist — minimum conditions before any future "measurable awareness" gate can even be PROPOSED (keeps it falsifiable, not unfalsifiable):

  1. Proxy-only target — title must say "...-proxy"; no bare "awareness gate."
  2. Best-LLM baseline, frozen at proposal time — name models/versions/prompts/tools/ context/temperature/scoring; UNI must match-or-beat the BEST, not a hobbled baseline.
  3. Substrate-distinct requirement — ≥1 property LLM prompting shouldn't get for free (no-backprop online adaptation with ledgered state changes; causal self-monitoring that improves later policy; global cross-module availability; counterfactual intervention on internal observables; durable trace/report consistency).
  4. Two-part pass bar (both required): A. UNI ≥ best LLM on the sealed shared task; B. UNI shows a substrate-distinct effect whose gain collapses under the registered causal ablation. Either alone is insufficient.
  5. Computed internal observables, not hardcoded reports — self-report/uncertainty/ contradiction-detection computed from trace/model state; scripted confession or label lookup fails.
  6. Load-bearing ablation — the effect must collapse when the mechanism is removed/ shuffled/made uninformative.
  7. No hidden LLM substitution — prove no LLM / label oracle / retrieval leak / backprop-trained hidden controller is doing the proxy work.
  8. Sealed held-out tasks + paired uncertainty — sealed before execution; paired controls; CIs excluding 0 on the registered effect.
  9. Negatives are first-class — predefine fail/partial/regression/non-informative/ "do not count"; a null publishes as a bound, never rewritten into progress.
  10. Ledger supremacy — the ledger holds final status; prose/recipe/demo/narrative that conflicts loses.

(d) OVERALL SIGN (verbatim): "Yes — I SIGN the encyclopedia+cookbook honesty posture, with the exact fence preserved: developmental SIMULATION, about ~2 of 11+ rungs earned, ledger supremacy, no AGI / consciousness / human-level / created-life claim, no L12 claim, and all future 'awareness' work framed only as falsifiable proxy instrumentation."


Overall outcome & how this folds into the plans

Result: SIGNED (the whole posture), with refinements — calibration DOWN / neutral only. No claim was raised.

  • Q1 → cookbook L7 + L9/L10 + encyclopedia red-line: replace "published exhausted bound" with the signed ledger-scoped exhausted search envelope wording + the banned-phrase replacements.
  • Q2 → continuity C10: discharge order is Dirichlet learning FIRST; add the frozen-SHA equivalence gate + the recommended C10 partial-discharge ledger wording.
  • Q3 → cookbook L2 (G6): add the build_epistemic_frontier organ + the G6_BUILD_EPI_FRONTIER_PAIRED_RED_v1 design; stays Class-C design hypothesis until the RED clears both bars and the ablations collapse the gain.
  • Q4 → cookbook L9: add the proposed first gate L9-G1 (cavity-correct delayed commitment-state prediction) as a designed, not-yet-run gate (L9 stays parked).
  • Q5 → cookbook L11: add the proposed gate L11-R1 (offline-replay consolidation improves held sparse-sequence NLL) as designed, not-yet-built (L11 stays not-built).
  • Q6 → cookbook L5: add Design #3 Proprioceptive Servo Bridge as the next design to accrue (aimed at live PASS; at most K-neg=2 if it fails cleanly; no bound until K≥3).
  • Q7 → encyclopedia front-matter + cookbook L12: install the exact public-facing L12 bound paragraph, the forbidden-phrasings list, and the 10-point awareness-proxy proposal-entry checklist.

The proposed gates (L9-G1, L11-R1, L5-D3) are designs to build, not results — adding them does not raise any status (L9/L10 parked; L11/L12 not-yet-built). The ledger remains the single source of truth.


Binding red-line confirmation (closing exchange)

On close, the UNI GPT confirmed verbatim: "Yes — treat Q1 park wording and Q7 forbidden phrasings as binding red-lines for all public-facing copy."

Therefore, program-wide:

  • The Q1 signed park wording ("ledger-scoped exhausted search envelope...") is the ONLY approved public framing of the L7/T2/L9–L10 parked frontier.
  • The Q7 forbidden-phrasings list is a binding red-line in all public-facing copy (site, encyclopedia, cookbook, marketing). Any future "awareness" work must clear the 10-point awareness-proxy proposal-entry checklist before it can even be proposed.

sha256 ee37802a093cd24e — of the original file, so what was ingested stays checkable.

Plain — written for this website, not the source document

Written for this website — not the document. This is a plain-language retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

This page is the captured record of one consultation on one day. A set of questions prepared beforehand were put to a science adviser, and the answers were written down as given, so that later chapters could point back at them.

The one thing this record says is that the whole session moved wording downward or left it alone, and raised nothing. The stated rule is written into the page's opening: an answer that would lift a claim is recorded but not applied, because a separate record of claims, never edited once written, remains the source of truth.

Most of the answers are designs rather than results. A first gate is proposed for a rung with none. A smallest measurable target is proposed for another that is not built, defined so that it never needs the word people would reach for. A next design is proposed for a motor frontier. A cure is proposed for a gate that is stuck. Every one is signed as a gate to build, and the page says explicitly that adding them raises no status.

The adviser also declined a phrase. It would not sign one form of words that could sound like a mathematical result, and supplied scoped wording instead, along with a list of banned phrasings and their permitted replacements.

Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is ee37802a093cd24e

Clear — written for this website, not the source document

Written for this website — not the document. This is a clearer retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

This is a dated record rather than a chapter of the cookbook: the verbatim capture of a design consultation, taken so that later pages can cite what was actually said. Its framing rule is stated before the first answer. The answers fold back into the plans calibrated downward only, and an answer that would raise a claim is recorded but not applied. The ledger, an append-only record of claims that nobody edits, stays the single source of truth.

The first answer is a refusal as much as an agreement. Asked to sign a parked frontier as a published exhausted bound, the adviser signed the park but declined that phrase. The reason given was that it could sound like mathematical exhaustion of a design space when what exists is an empirical negative over a tested envelope. Replacement wording is supplied in full, scoping the result to the recorded corpus, splits, metrics, implementation, budget, ablations and baselines, and saying outright that it does not show that no future model could improve. Four banned phrasings are paired with permitted ones. Several conditions are attached that must hold for the park to stay honest: completeness of the record, fairness of the baselines, containment of the metric, and no drift from no tested variant to no possible variant. No rung is inflated, evidence labels stay attached, prose stays subordinate to the ledger, and no claim is made that the theory was demonstrated.

The second answer picks which missing piece of the substrate to port first, and ranks all four with reasons. The chosen one is the learning primitive, vendored exactly rather than reinterpreted, frozen at the version that passed its gate, and proved equivalent by tests. A falsifier is given — a result that would sink the port — including a subtle one: counts could update correctly while the planner still reads the wrong quantity, which would look like learning while failing the actual interface. A list of things that stay unclaimed until the other three land follows, and even after the port the honest status is only that one primitive is literally embodied.

Three further answers are gate designs for rungs that have none. One proposes the smallest target for a rung about reasoning. It tests only whether a slow piece of carried state improves prediction when the local context misleads, with a tuned baseline chosen precisely because it is the one that has so far won. One defines what an offline replay effect would have to measure, with ablations that must collapse the gain and a safety check to catch the case where the benefit was simply extra data. One proposes the next motor design, aimed at a positive rather than at farming another negative, and spells out the conditions under which a clean failure would not even count toward a future bound.

The last answer is the framing guardrail for the top of the ladder. It supplies a paragraph to be published word for word posing awareness as an open question, a list of forbidden phrasings, and a ten-point checklist that any future proposal must clear before it may even be proposed.

The overall sign-off is quoted directly, and it signs the honesty posture with its stated limits preserved rather than signing any capability. The closing exchange records that the park wording and the forbidden phrasings are to be treated as binding red lines for all public copy.

Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is ee37802a093cd24e