UNI-FLAGELLUM · point-in-time snapshot

Where the model actually stands

Measured 2026-08-19 against the committed artifacts, not the prose. Every figure below carries its source file. This page is a snapshot — it will age; the re-measure commands are at the foot. Overall machine verdict: PARTIAL_PARITY_ONLY science-gates-report.json:5

Adverse first

None of the three model stacks has earned biological parity, and the deepest gap is one no code can close.

1 · The behaviour we must replicate — the yardstick

From the Wikipedia flagellum article and the model's own primary anchors (Antani 2021, Wadhwa 2022, Lo 2018, Mattingly & Tu 2026). Our model is an explicit teaching reduction of this, not a molecular replica. SCIENCE.md:24-42

Machine

  • Stator MotA₅:MotB₂, 11–16 units, proton-driven; more recruited under load
  • Rotor C-ring switch FliG/FliM/FliN
  • Rings L·P·MS · hook · flagellin filament (~20 nm, 11 protofilaments)
  • Stator↔rotor gear ratio ≈6.2

Physics

  • Energy = proton-motive force (Na⁺ in Vibrio)
  • Loaded 200–1000 rpm; unloaded 6,000–100,000 rpm
  • Stall torque ∝ stators × PMF; speed falls with load
  • Low-Reynolds regime

Behaviour

  • CCW = run (bundle, forward); CW = tumble (unbundle)
  • CheY-P → C-ring biases CW → tumbling
  • Run-and-tumble = biased random walk = chemotaxis
  • Adaptation via receptor methylation

The three species the truth contract keeps separate: E. coli (behavioural — run/tumble, PMF), Salmonella and Bacillus (structural — switch complex, protofilaments). They are evidence sources, never one measured specimen. CLAUDE.md:207-209 · lib/walkthrough.js:28,43,58

1·5 · Measured this run — 2026-08-19, actually executed on this machine

① JS21 / 21

node v25.0.0 · node --test

All invariants hold: VFE identity, log-odds identity, Markov boundary, determinism. A passing test confirms the adverse held-out comparison is surfaced, not suppressed.

② Python573 / 577

pytest 8.3.3 · numpy 2.3.5 · scipy 1.16.3

Settles the count: the docs said 71 and 533 — both wrong; measured is 577 collected / 573 pass / 3 skip / 1 xfail. D5 firewall respected; the lognormal-beats-mixture result reproduces green.

③ Elixir2 fail

1.19.5 / OTP 28 · mix test · 1049 tests

Engine passes; the 2 reds are stale hardcoded ledger tallies (tests expect PASS 93/PARTIAL 4, live ledger is 94/5) — an operator-gated pin, not an engine bug. Benchmark: all six baselines incl. Random survive to 300 ticks — no survival separation yet. ledger_schema_conformance_test.exs:145 · verdict_partial_names_subclaim_test.exs:101

2 · The three model stacks

① JS agentcorrectness only

lib/uni-motor.js · runs in the browser lab

World → Markov boundary → discrete active-inference agent. Hidden gradient state {falling, flat, rising}; exact categorical posterior; VFE = surprise (KL≡0 by construction); EFE policy over RUN / TUMBLE, softmax γ=4.

Runs: tick-by-tick in the UI. Gated by: algebraic invariants + cross-engine determinism only.

Adverse: never scored against biological data — un-gated against data by design. EFE uses a non-canonical form that double-weights ambiguity (code & doc agree; no oracle checks it). uni-motor.js:328

② Python competitionNOT_ESTABLISHED

hierarchical-aif/ · fitted & scored on Wadhwa-2022

The one place models are literally run and scored on real held-out E. coli dwell data: 80 train / 19 holdout motors, motor-equal NLPD, paired motor-cluster bootstrap. compare.py · seed 20260717

Candidate F_MOTOR_STACK · control M3 two-timescale mixture · adversaries M0/M1/M2/M5/M8.

Adverse: lognormal (adversary) wins; F-side reproduces the incumbent M7 to 2.5e-7 nats — a re-derivation, not a new model; buys ~nothing over its own τ→0 Weibull limit. REPORT.md:20-32,113-121

③ Elixir colony@scaffold

sibling repo UNI.Minecraft · SP.Sim + SP.Runtime

Pure open-ended world + a live Minecraft colony that breeds brains on death. The pure benchmark runs offline today (mix test, seeds 101-106). A trained lineage exists on disk. runs/colony/kin-8.bin

Graduation gate forage-pureworld-graduation: PENDING — the twin-lineage harness (Twin A trained vs Twin B untrained), the N≥8 seed sweep, and the receipt writer are unbuilt. gates.ndjson:5

Adverse: nearest predecessor nursery-fenced-red-stocked FAIL (a bot starved — geometry, not policy). Building the harness is a BUILD, operator-gated, needs live MC + ~64 colony-hours off-broadcast.

3 · The parity ladder — how far up the evidence licenses

P8 is conjunctive: one non-PASS required rung ⇒ FULL_PARITY = false. The ladder is defined generically in CLAUDE.md; per-level status from the H-AIF receipt map. BIOLOGICAL-PARITY-RECEIPT-MAP.md · HIERARCHICAL-AIF-GATE-TO-EXISTING-P-LADDER-MAP.md

P0Computational integrity — frozen SHA-256 baseline; all audit hashes reproduceHOLDS
P1Equation / implementation — first-passage math PASS; but G03 & G05 FAIL live insideHOLDS*
P2Observational — 11 studies, ≥409 motors, source/aggregate level; raw MAT tier absentLIMITED
P3Held-out predictive — performed & scored, but the mechanism LOSES to a lognormalACTIVITY ONLY
▲ the software wall — nothing below this line can be crossed by code in this repo ▲
P4Transfer — 0 commensurate cross-lab transfer tests · FIRST UNSATISFIED RUNGNOT_ESTABLISHED
P5Intervention — no discriminating intervention; "does a bacterium infer?" unanswerable hereNOT_ESTABLISHED
P6Structural / mechanistic — scoped; lattice L=13 is not resolved molecular geometry (X11 FAIL)SCOPED / FAIL
P7Independent replication — independent laboratories: 0BLOCKED_EXTERNAL
P8Full verdict — conjunction of all required rungsFALSE

Honesty flag: the docs call P4 the first unsatisfied rung by treating P3 as "satisfied-as-activity". Under a strict "first rung without a PASS" reading it is P3, because the held-out mechanistic-advantage gate G06 FAILs. Both are stated so you can rule. G06 advantage 0.0583 nat/interval, 95% CI [-0.0158, 0.1326] crosses 0

4 · The gates, as the artifacts report them

Science gates G00–G13 — single-study

PASS 4 · FAIL 3 · SOURCE_ONLY 1 · NOT_ESTABLISHED 1 · BLOCKED_EXTERNAL 5 science-gates-report.json:5-19

G00 sourceG01 boundaryG02 math G03 artifactG04 censoredG05 recovery G06 held-outG07 H-stateG08 load/torque G09 switchG10 AIF identityG11 live G12 replicationG13 physical

Cross-study gates X01–X16

PASS 8 · FAIL 3 · NOT_ESTABLISHED 2 · BLOCKED_EXTERNAL 3 cross-study-parity-report.json:776-789

X01X02X03X04 X05X06 latticeX07X08 X09X10 transferX11 structuralX12 AIF X13 liveX14 wet-labX15 printX16 full

5 · The held-out leaderboard — where the mechanism actually loses

Motor-equal NLPD on 19 holdout motors (lower = better). The experimental unit is the motor; every contrast interval crosses zero, so nothing is "established" and nothing is called "equivalent". F-SIDE-MOTOR-STACK-SCORING-REPORT.md:70-78

RankModelRoleNLPD
1M2 lognormaladversary3.4093← the simple curve wins
2M8 empirical KDEadversary (no mechanism)3.4225
3F_MOTOR_STACKcandidate3.4327= re-derivation of M7 (2.5e-7 nats)
4M1 Weibulladversary (τ→0 limit)3.4333candidate buys ~nothing over this
5M3 two-timescale mixturecontrol / "current design"3.4343
6M5 gammaadversary3.4637
7M0 exponentialmemoryless null3.5480only model the candidate clearly beats

6 · "The three" — genuinely underdetermined

The repo fixes no canonical "the three." The reading where every verb of your phrase lands — "model the three fully / run them / work the gates" — is the three runnable stacks above (JS / Python / Elixir). That's my working assumption, at low confidence. The live runnable set is the same work under the strongest alternatives, so I've started running all three regardless.

Candidate meaningVerdict
Three stacks — JS / Python / Elixirbest fit for "run them"; unattested as a named trio
Three roles — target / control / adversarya real framing, but there are 5 adversaries and you can't "run" a target hypothesis
Three duration models — mixture / lognormal / nullthe repo's only literal "three models" is M0/M1/M2 (exp/Weibull/lognormal) — excludes the mixture
Three species — E. coli / Salmonella / Bacillusthe stable triad, but evidence sources, not runnable models
Three fleet bodiesfalsified — the fleet is four (Door/Control-Plane/Gaia/HUD)

This one is yours: which "three" did you mean — and if it's the stacks, do you want me to spend the operator-gated Elixir twin-harness build (~64 colony-hours, off-broadcast)? I won't sink that on a guess.

7 · The road to "fully replicating the behaviour"

What code CAN move (in-repo)

  • Close the doc-vs-artifact drift cluster (76/16→43/11; stale H-AIF-GATES/README; measured test counts)
  • Run all three stacks & record real receipts (in flight now)
  • Score the JS agent against a commensurate observable — or honestly report that none exists
  • Add an EFE-decomposition oracle; re-run the Python competition against M0–M8
  • Build the Elixir twin-harness so the graduation gate can run (operator-gated)
  • Addressable without wet-lab data: G03 (needs a tagged source-parameter artifact) & G07 (needs an H-well classifier)

What no code can move — the wall

  • P4 transfer — needs an independent cross-lab cohort
  • P5 intervention — needs a discriminating biological intervention (this is what "does a bacterium infer?" requires)
  • P7 replication — needs an independent wet lab
  • 8 gates carry softwareCannotSubstitute:true — G08/G09/G11/G12/G13, X13/X14/X15
  • The D5 holdout mark channel was irreversibly burned (2026-07-21); re-earnable only on independent data

The "world where we fully replicate the behaviour in inference" is a target world defined by external receipts we do not yet hold — not a software finish line.

7·5 · Findings from this session's deeper lanes

Drift closed11 fixes

11 doc-vs-artifact lies corrected across 6 files, verified against the artifacts: SCIENCE-GATES 76/16→43/11; H-AIF-GATES G5/G6/G7 stale statuses; README & CURRENT-STATE test counts → measured 573/577; hierarchy.py Lmotor label; scope-ruling LANE B. Uncommitted.

Not tri-speciesE. coli only

The runnable model is E. coli only. Salmonella (2 assets) and Bacillus (1) are structural illustration — zero runnable dynamics, zero scored gates — and the separation is genuinely code-enforced (walkthrough.test.mjs pins each species). One conflation risk for your eye: the E. coli-scoped structural-state-map imports the Salmonella cryo-EM paper under a different citation label without restating the source species.

EFE self-check2.6% flip

The agent's expected-free-energy double-weights ambiguity. Sweeping all 5,151 belief states, that changes the RUN/TUMBLE decision on 134 (2.6%) — a thin near-tie band (margin ≤ 0.034 nat) where the gradient is believed flat: there the extra term tips the agent toward exploratory TUMBLE. Not random — it amplifies information-seeking exactly in the flat regime (the very AIF signature the P5 door tests). Whether intended is the operator's ruling; no gate checks it. scratchpad/efe_probe.mjs vs uni-motor.js:309-328

8 · Provenance & re-measure

This page is a snapshot dated 2026-08-19. To re-measure rather than trust it:

QuestionCommand
JS model invariantsnode --test tests/model.test.mjs
Science gates (deterministic)npm run science:run && npm run science:verify
Cross-study paritynpm run cross-study:run && npm run cross-study:verify
Python motor-stack suitepython -m pytest tests/motor_stack_aif -q
Elixir pure world (sibling repo)cd ../UNI.Minecraft && mix test
Are the trees clean?git status -sb

Authoritative artifacts: experiments/results/science-gates-report.json, experiments/results/cross-study-parity-report.json, hierarchical-aif/reports/F-SIDE-MOTOR-STACK-SCORING-REPORT.md, hierarchical-aif/ledgers/HIERARCHICAL-AIF-GATE-TO-EXISTING-P-LADDER-MAP.md, UNI.Minecraft/evidence/gates.ndjson.

Snapshot generated from the 6-agent model-design audit · UNI-FLAGELLUM · uncommitted, for review.