Measured 2026-08-19 against the committed artifacts, not the prose. Every figure below carries its source file. This page is a snapshot — it will age; the re-measure commands are at the foot. Overall machine verdict: PARTIAL_PARITY_ONLY science-gates-report.json:5
None of the three model stacks has earned biological parity, and the deepest gap is one no code can close.
@scaffold. UNI.Minecraft/runs/pureworld_qa_gate.exs:28,39-40SCIENCE-GATES.md:58 says G04 has 76/16 censored intervals; the artifact says 43/11 (known, unrepaired; may be an eligible-state subset — verify before "fixing"). science-gates-report.json:248-249From the Wikipedia flagellum article and the model's own primary anchors (Antani 2021, Wadhwa 2022, Lo 2018, Mattingly & Tu 2026). Our model is an explicit teaching reduction of this, not a molecular replica. SCIENCE.md:24-42
The three species the truth contract keeps separate: E. coli (behavioural — run/tumble, PMF), Salmonella and Bacillus (structural — switch complex, protofilaments). They are evidence sources, never one measured specimen. CLAUDE.md:207-209 · lib/walkthrough.js:28,43,58
node v25.0.0 · node --test
All invariants hold: VFE identity, log-odds identity, Markov boundary, determinism. A passing test confirms the adverse held-out comparison is surfaced, not suppressed.
pytest 8.3.3 · numpy 2.3.5 · scipy 1.16.3
Settles the count: the docs said 71 and 533 — both wrong; measured is 577 collected / 573 pass / 3 skip / 1 xfail. D5 firewall respected; the lognormal-beats-mixture result reproduces green.
1.19.5 / OTP 28 · mix test · 1049 tests
Engine passes; the 2 reds are stale hardcoded ledger tallies (tests expect PASS 93/PARTIAL 4, live ledger is 94/5) — an operator-gated pin, not an engine bug. Benchmark: all six baselines incl. Random survive to 300 ticks — no survival separation yet. ledger_schema_conformance_test.exs:145 · verdict_partial_names_subclaim_test.exs:101
lib/uni-motor.js · runs in the browser lab
World → Markov boundary → discrete active-inference agent. Hidden gradient state {falling, flat, rising}; exact categorical posterior; VFE = surprise (KL≡0 by construction); EFE policy over RUN / TUMBLE, softmax γ=4.
Runs: tick-by-tick in the UI. Gated by: algebraic invariants + cross-engine determinism only.
Adverse: never scored against biological data — un-gated against data by design. EFE uses a non-canonical form that double-weights ambiguity (code & doc agree; no oracle checks it). uni-motor.js:328
hierarchical-aif/ · fitted & scored on Wadhwa-2022
The one place models are literally run and scored on real held-out E. coli dwell data: 80 train / 19 holdout motors, motor-equal NLPD, paired motor-cluster bootstrap. compare.py · seed 20260717
Candidate F_MOTOR_STACK · control M3 two-timescale mixture · adversaries M0/M1/M2/M5/M8.
Adverse: lognormal (adversary) wins; F-side reproduces the incumbent M7 to 2.5e-7 nats — a re-derivation, not a new model; buys ~nothing over its own τ→0 Weibull limit. REPORT.md:20-32,113-121
sibling repo UNI.Minecraft · SP.Sim + SP.Runtime
Pure open-ended world + a live Minecraft colony that breeds brains on death. The pure benchmark runs offline today (mix test, seeds 101-106). A trained lineage exists on disk. runs/colony/kin-8.bin
Graduation gate forage-pureworld-graduation: PENDING — the twin-lineage harness (Twin A trained vs Twin B untrained), the N≥8 seed sweep, and the receipt writer are unbuilt. gates.ndjson:5
Adverse: nearest predecessor nursery-fenced-red-stocked FAIL (a bot starved — geometry, not policy). Building the harness is a BUILD, operator-gated, needs live MC + ~64 colony-hours off-broadcast.
P8 is conjunctive: one non-PASS required rung ⇒ FULL_PARITY = false. The ladder is defined generically in CLAUDE.md; per-level status from the H-AIF receipt map. BIOLOGICAL-PARITY-RECEIPT-MAP.md · HIERARCHICAL-AIF-GATE-TO-EXISTING-P-LADDER-MAP.md
Honesty flag: the docs call P4 the first unsatisfied rung by treating P3 as "satisfied-as-activity". Under a strict "first rung without a PASS" reading it is P3, because the held-out mechanistic-advantage gate G06 FAILs. Both are stated so you can rule. G06 advantage 0.0583 nat/interval, 95% CI [-0.0158, 0.1326] crosses 0
PASS 4 · FAIL 3 · SOURCE_ONLY 1 · NOT_ESTABLISHED 1 · BLOCKED_EXTERNAL 5 science-gates-report.json:5-19
PASS 8 · FAIL 3 · NOT_ESTABLISHED 2 · BLOCKED_EXTERNAL 3 cross-study-parity-report.json:776-789
Motor-equal NLPD on 19 holdout motors (lower = better). The experimental unit is the motor; every contrast interval crosses zero, so nothing is "established" and nothing is called "equivalent". F-SIDE-MOTOR-STACK-SCORING-REPORT.md:70-78
| Rank | Model | Role | NLPD | |
|---|---|---|---|---|
| 1 | M2 lognormal | adversary | 3.4093 | ← the simple curve wins |
| 2 | M8 empirical KDE | adversary (no mechanism) | 3.4225 | |
| 3 | F_MOTOR_STACK | candidate | 3.4327 | = re-derivation of M7 (2.5e-7 nats) |
| 4 | M1 Weibull | adversary (τ→0 limit) | 3.4333 | candidate buys ~nothing over this |
| 5 | M3 two-timescale mixture | control / "current design" | 3.4343 | |
| 6 | M5 gamma | adversary | 3.4637 | |
| 7 | M0 exponential | memoryless null | 3.5480 | only model the candidate clearly beats |
The repo fixes no canonical "the three." The reading where every verb of your phrase lands — "model the three fully / run them / work the gates" — is the three runnable stacks above (JS / Python / Elixir). That's my working assumption, at low confidence. The live runnable set is the same work under the strongest alternatives, so I've started running all three regardless.
| Candidate meaning | Verdict |
|---|---|
| Three stacks — JS / Python / Elixir | best fit for "run them"; unattested as a named trio |
| Three roles — target / control / adversary | a real framing, but there are 5 adversaries and you can't "run" a target hypothesis |
| Three duration models — mixture / lognormal / null | the repo's only literal "three models" is M0/M1/M2 (exp/Weibull/lognormal) — excludes the mixture |
| Three species — E. coli / Salmonella / Bacillus | the stable triad, but evidence sources, not runnable models |
| Three fleet bodies | falsified — the fleet is four (Door/Control-Plane/Gaia/HUD) |
This one is yours: which "three" did you mean — and if it's the stacks, do you want me to spend the operator-gated Elixir twin-harness build (~64 colony-hours, off-broadcast)? I won't sink that on a guess.
The "world where we fully replicate the behaviour in inference" is a target world defined by external receipts we do not yet hold — not a software finish line.
11 doc-vs-artifact lies corrected across 6 files, verified against the artifacts: SCIENCE-GATES 76/16→43/11; H-AIF-GATES G5/G6/G7 stale statuses; README & CURRENT-STATE test counts → measured 573/577; hierarchy.py Lmotor label; scope-ruling LANE B. Uncommitted.
The runnable model is E. coli only. Salmonella (2 assets) and Bacillus (1) are structural illustration — zero runnable dynamics, zero scored gates — and the separation is genuinely code-enforced (walkthrough.test.mjs pins each species). One conflation risk for your eye: the E. coli-scoped structural-state-map imports the Salmonella cryo-EM paper under a different citation label without restating the source species.
The agent's expected-free-energy double-weights ambiguity. Sweeping all 5,151 belief states, that changes the RUN/TUMBLE decision on 134 (2.6%) — a thin near-tie band (margin ≤ 0.034 nat) where the gradient is believed flat: there the extra term tips the agent toward exploratory TUMBLE. Not random — it amplifies information-seeking exactly in the flat regime (the very AIF signature the P5 door tests). Whether intended is the operator's ruling; no gate checks it. scratchpad/efe_probe.mjs vs uni-motor.js:309-328
This page is a snapshot dated 2026-08-19. To re-measure rather than trust it:
| Question | Command |
|---|---|
| JS model invariants | node --test tests/model.test.mjs |
| Science gates (deterministic) | npm run science:run && npm run science:verify |
| Cross-study parity | npm run cross-study:run && npm run cross-study:verify |
| Python motor-stack suite | python -m pytest tests/motor_stack_aif -q |
| Elixir pure world (sibling repo) | cd ../UNI.Minecraft && mix test |
| Are the trees clean? | git status -sb |
Authoritative artifacts: experiments/results/science-gates-report.json, experiments/results/cross-study-parity-report.json, hierarchical-aif/reports/F-SIDE-MOTOR-STACK-SCORING-REPORT.md, hierarchical-aif/ledgers/HIERARCHICAL-AIF-GATE-TO-EXISTING-P-LADDER-MAP.md, UNI.Minecraft/evidence/gates.ndjson.
Snapshot generated from the 6-agent model-design audit · UNI-FLAGELLUM · uncommitted, for review.