dashboardTHE CLASSROOMstate snapshotexternal doors
The biological programme, end to end: where the data comes from, what the models are, how they are scored, and which gates stand. Everything below was read from the code and the machine-readable result files, and the runnable parts were executed to confirm they still reproduce.
On the only real held-out E. coli data in this estate, a plain lognormal out-predicts our two-timescale mechanism and every mechanistic candidate. Every contrast interval crosses zero, so the honest word is NOT_ESTABLISHED — not "refuted", and never "equivalent".
Full biological parity is FALSE. Cross-lab transfer tests: 0. Discriminating interventions: 0. Independent laboratories: 0.
Not a guess. The released pipeline scores our model against exactly three named rivals, and the object holding them has exactly three keys.
The three are the rivals our model must beat. The two-timescale mixture is the fourth because it is ours — you do not beat yourself.
Motor-equal negative log predictive density on the seconds scale, 19 held-out motors, lower is better. Our model places 4th of 4.
All of it runs. Nothing here is blocked on missing code. What is missing is evidence, not runnability.
577 tests. The full F-side competition re-runs end to end in ~21 s and reproduces every frozen number to the last digit. Verified today.
Reproduces every frozen number in-memory. Verified today.
Deterministic: running twice must produce byte-identical report and audit files, or the run is rejected.
PASS 4 FAIL 3 SOURCE_ONLY 1 NOT_ESTABLISHED 1 BLOCKED_EXTERNAL 5
G03 the published parameter vector does not reproduce the article's own figure arrays · G05 only 2 of 3 parameter-recovery replicates pass · G06 the mechanism does not beat memoryless out-of-sample (interval crosses zero)
PASS 8 FAIL 3 NOT_ESTABLISHED 2 BLOCKED_EXTERNAL 3
X06 the 13-site lattice gives incompatible interaction strengths from distributions vs moments · X11 the lattice is not resolved molecular geometry · X16 full biological parity = FAIL
| rung | what it demands | status |
|---|---|---|
| P0 | computational integrity — hashes, determinism | HOLDS |
| P1 | equations match implementation | HOLDS* |
| P2 | observational — real recorded measurements | LIMITED |
| P3 | held-out prediction | DONE — AND THE MECHANISM LOST |
| P4 | transfer — parameters frozen on one lab predict another's raw data | FIRST UNSATISFIED · 0 tests |
| P5 | intervention — does a bacterium actually infer? | 0 interventions |
| P6 | structural / mechanistic identity | SCOPED · X11 FAIL |
| P7 | independent replication | 0 laboratories |
| P8 | the conjunction of all of the above | FALSE |
All four models already exist and already run. So "fully" cannot mean "write them" — it means strengthen the comparison until the result is decisive rather than inconclusive. Concretely, three things, in order:
| # | act | why it is the honest next step |
|---|---|---|
| 1 | Re-run the competition to a new output path and confirm the frozen numbers still reproduce, without touching the frozen record | establishes the baseline is live rather than quoted — everything I told you today was copied from July |
| 2 | Make the comparison decisive: every contrast currently crosses zero at 19 holdout motors. Compute the motor count at which a 0.042-nat difference resolves, from a real-data power analysis — never chosen after seeing an interval | "underpowered" is not "equivalent". Right now we cannot tell whether the mechanism loses or we cannot see |
| 3 | Fix the two verified defects in the released agent: its free energy is vacuous (the identity it tests is true by construction) and its expected free energy double-weights ambiguity | both confirmed numerically today; the second changes which action it picks |
Read from the code and the machine-readable results; runnable parts executed to confirm reproduction. Numbers are the frozen competition's, which this session did not re-run — that is act 1 above.