UNI Universal Natural Intelligence

Wiki · Evidence & Verdicts

RED pre-registration — depth-red-b

Evidence & Verdicts · docs/receipts/red_preregistration_depth_red_b.md @ 44baf03d5041 (gen2-runtime) — opens the published snapshot ac338733bbba

How to read this page

Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.

Eighty-seven dated pages: receipts, pre-registrations, handoffs, validation records and review verdicts. A receipt is written at the moment a piece of work was checked. It names what was claimed, the commit and the seed, what was actually run, and the outcome in one of a small set of controlled words. Then it names what the work did not achieve. That last part is what makes it a receipt rather than an announcement. A pre-registration is the same discipline run in advance: the conditions that would count as a pass and the conditions that would falsify the claim are written down before the run, so neither can be adjusted once the numbers arrive.

That is why so many small dated stubs are an audit trail rather than noise. No one of them is meant to be a good read. The value is in the sequence and in the dates, because you can watch a prediction be registered, then the run happen, then the verdict land — sometimes against the prediction. Pages here record a falsified result, a rejected fix, a retracted overclaim, and a green receipt that turned out not to be reproducible from the commit that carried it. A record that carried only successes would be worth a good deal less than this one.

A gentle way in is to read a pre-registration first, so the shape becomes familiar, then a result page, then one of the corrections. This section sits off the main navigation on purpose: it is the record you check the rest of the site against, not the place to begin.

What it is not: documentation, and not a summary. Nothing here has been tidied in hindsight. Every entry reads as of its date, a later entry may overturn an earlier one, and the presence of a page is not a claim that its result stood.

Your browser cannot switch reading levels, so the document itself is shown.

Precise — the source document

This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.


verdict: WITHHELD evidence_class: pending

RED pre-registration — depth-red-b

  • Gate name: depth-red-b
  • Phase: Phase 2b (Sensorium)
  • Pre-registered: 2026-07-13
  • Runner: runs/depth_red.exs
  • Related: docs/specs/sensorium.md:5-40

Motivation

docs/specs/sensorium.md pre-registers RED-B for the :depth factor with init_a: :diagonal — the identifiability gate. The RED must run AFTER RED-A has a verdict (one-cure-at-a-time; RED-A is the vision factor).

PASS condition

Under init_a: :diagonal for the :depth factor, posterior separates the depth prior from the vision prior on the pre-registered ablation set: the posterior over depth-states is distinguishable from the posterior over vision-states at every tick of the diagnostic window.

FALSIFIES condition

  • Depth-factor identifiability collapses (posteriors indistinguishable, KL(depth‖vision) < ε on the ablation set), OR
  • default_genome byte-identity breaks with :depth absent (test/sp/brain/decider_byte_identity_test.exs fails).

Protocol

  1. Genome: depth_lineage/0 (new lineage, absent from default/0). Coupling 0.0 by default.
  2. depth.a seeded with init_a: :diagonal; vision.a unchanged.
  3. Ablation set: pre-registered N=100 scene tuples with independent depth/vision ground truth.
  4. Run inference for K=1024 ticks. Sample posteriors every 128 ticks.
  5. Verdict:
    • PASS: KL(depth ‖ vision) > threshold (pre-registered ε = 0.5 nats) on all sampled ticks.
    • PARTIAL: PASS on a majority of sampled ticks.
    • FAIL: any decider byte-identity test failure, OR posteriors indistinguishable.

sha256 b27be9338a9c83df — of the original file, so what was ingested stays checkable.

Plain — written for this website, not the source document

Written for this website — not the document. This is a plain-language retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

A promise written before a test was run, and not a result. It says what would be tried, what would count as a pass, and what would count as a refutation, and then it stops, because at the time of writing there was no verdict. The question it sets up is whether two things a model is meant to work out separately can really be told apart, or whether they collapse into one. If they collapse, the page says plainly, in advance, that this counts as a failure.

Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is b27be9338a9c83df

Clear — written for this website, not the source document

Written for this website — not the document. This is a clearer retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

No result has been recorded here at all. It is a pre-registration: the test, the pass condition and the result that would show it wrong, all committed to paper before the run, with the verdict withheld.

The header block names the gate, the phase of work it belongs to, the date it was registered, the script that would run it, and the specification it comes from. One ordering rule is stated up front. This test must wait until an earlier, related test has a verdict, so that only one change is under examination at a time.

The pass condition asks that the model's belief about one factor stay distinguishable from its belief about another at every sampled point of a diagnostic window. What would show it wrong is the mirror of that. If the two beliefs become indistinguishable, or if an unrelated default behaviour stops being byte-identical, the test fails. Then comes the protocol, which fixes the details so they cannot drift later: which genome, how it is seeded, the scene set to run against, how long to run, and how often to sample. It ends with a small table turning the measurement into pass, partial or fail.

Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is b27be9338a9c83df