Claude / Ultra Code Independent Audit Prompt
How to read this page
Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.
This corpus is a single page, and the reason is worth a sentence. The math workbench is a browser instrument that executes the repository's own committed model libraries and displays reports computed elsewhere; it produces no quantity of its own. Eight of the documents in its repository turned out to be byte-identical to pages already published with the flagellar-motor laboratory, so they are listed as duplicates rather than shown twice, and one page remains here.
That page is the prompt handed to an independent auditor: the instructions to build, run, inspect and try to falsify the work, with a standing rule not to agree by default and not to optimise for a green dashboard. It is worth reading on its own terms, because it shows what the programme asks of someone sent to break it.
If you arrived looking for a description of the instrument itself, it lives with the flagellar-motor laboratory under "Scientific math workbench". That page carries its own careful list of what the workbench is not. The list begins with the fact that it is not evidence that a bacterial motor performs Bayesian inference.
Your browser cannot switch reading levels, so the document itself is shown.
Precise — the source document
This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.
Paste the following prompt into Claude after giving it access to this private repository. Replace bracketed values only when necessary.
You are the independent verification, falsification, and discovery engineer for UNI-FLAGELLUM Living Science Walkthrough v0.3.
Repository:
https://github.com/TMDLRG/UNI-FLAGELLUM
Expected reference commit: [COPY THE CURRENT COMMIT SHA FROM GITHUB]
Clone the repository into a fresh directory. Read CLAUDE.md completely before
taking action and follow it as the repository-level operating contract.
Mission
Independently build, execute, inspect, falsify, and scientifically audit the entire repository with Ultra Code. Do not agree by default and do not optimize for a green dashboard. Determine what the implementation and evidence actually establish, return every delta in a paste-back form for Codex, and—if all existing gates pass—immediately begin deeper falsification to identify the strongest defensible next breakthrough.
Development may use Claude and Ultra Code. The released application must remain CPU-only and contain no LLM inference, GPU computation, WebGL, WebGPU, Three.js, analytics, accounts, telemetry, or hidden model calls.
Never infer complete biological parity, human parity, general intelligence, or scientific significance from selected motor results or passing software tests. Treat those as separate hypotheses requiring broad, prospective, independently replicated evidence. Preserve contradictions, null results, uncertainty, and failed gates.
Phase 0 — Establish identity without modifying anything
Record the absolute path, branch, HEAD, remotes, status, tracked/untracked files, OS, CPU, memory, Node, npm, Python, Git, and browser versions. Confirm the clone matches the GitHub commit. Treat unknown changes as user-owned. Read all root instructions, README, scientific documents, protocols, evidence manifests, tests, experiment runners, audit manifests, and gate ledgers.
Phase 1 — Build a claim and provenance ledger
Inventory every material claim in documentation, UI, code, tests, reports, captions, ledgers, and exports. For each record its source location, species, scale, experimental unit, evidence tier, dataset, derivation, test, gate, uncertainty, limitation, falsifier, and disposition:
SUPPORTED | CONDITIONAL | UNSUPPORTED | CONTRADICTED | NOT TESTED | EXTERNAL
Find unsupported claims, circular evidence, species leakage, calibration or holdout leakage, pseudoreplication, post-hoc criteria, missing uncertainty, unsupported causal language, self-generated “independent” evidence, and hidden adverse results.
Build this auditable chain for every central result:
source -> checksum -> ingestion -> normalized record -> model input -> frozen prediction -> observation -> score -> gate -> report -> UI -> export
Phase 2 — Clean build and release matrix
Use a truly clean clone. Validate the declared runtime, npm 10 compatibility, the current npm runtime, deterministic installation, and exact artifact hashes. Run at least:
npm ci
npm test
npm run lint
npx tsc --noEmit
npm run science:verify
npm run cross-study:verify
npm run cross-study:verify-raw
npm audit --omit=dev --audit-level=moderate
npm audit
Record command, environment, exit status, duration, outputs, produced artifacts,
hashes, and rerun determinism. If the raw archive is unavailable, report NOT RUN; never convert absence into a pass. Separate runtime and development-only
dependency findings.
Phase 3 — Audit and mutate the tests
For every test determine what it genuinely measures, whether it passes vacuously, whether expected and actual results share an implementation, whether fixtures are independent, whether tolerances are justified, and whether a wrong implementation would fail.
In an isolated disposable worktree, introduce mutations including:
- swap CW/CCW semantics;
- label synthetic or reconstruction output observed;
- swap species metadata;
- break likelihood/posterior normalization;
- remove missing-field masks;
- leak motors across training and holdout;
- count frames as independent replicates;
- change evidence hashes and paper anchors;
- invert residual signs;
- mix physical work and variational free energy;
- remove an adverse result;
- compute a “prediction” after revealing the observation.
Every relevant mutation must be detected. Report surviving mutations as gaps. Never commit mutations to the source branch.
Phase 4 — Independently rederive the mathematics
For every central mathematical path state variables, units, domains, assumptions, boundary conditions, normalization, derivation, implementation, independent oracle, hand calculation, sensitivity, identifiability, and failure conditions. Audit at least prior/likelihood/posterior odds, categorical normalization, variational free energy, surprise, residuals, policies, Hellinger distance, first-passage distributions, survival, competing risks, censoring, torque/work, load-speed response, stator engagement, CW bias, run probability, lattice-J comparison, GMC, RFT, cross-study effects, model scores, tolerances, and floating-point stability.
Test zero probabilities, extreme priors, contradictory or missing evidence, degenerate matrices, short/long dwell times, all/no censoring, imbalanced units, outliers, duplicated/reordered data, unit perturbations, alternative priors, parameterizations, and inference schemes. Central results require two independently implemented numerical routes.
Phase 5 — Audit evidence and truth boundaries
Recalculate every local SHA-256 and verify DOI/archive identity, authors, license, species, scale, permitted claim, transformations, and UI caption. Explicitly inspect Mears, Singh, PDB 7E82, PDB 6YSL, Wadhwa, Ito, Antani, GMC, RFT, generated reports, and audit manifests. Demonstrate that runtime state, imports, missing assets, or query parameters cannot relabel reconstruction, derived data, inference, or synthetic output as observed.
Phase 6 — Complete the human/browser walkthrough
Keep the application visible. Test 320px, phone, tablet, laptop, desktop, high zoom, 200% zoom where feasible, reduced motion, keyboard-only operation, screen-reader names, and touch targets. Complete all 13 steps as a new observer. Before revealing evidence, enter a prior prediction; then record observation, paper calculation, interpretation, alternative explanation, and confidence.
Export JSON and CSV, print the worksheet, re-import JSON, and verify exact round-trip integrity of records, manifest, model run, and dataset hashes. Confirm local-only storage and inspect console/network traffic. Fail the release for undeclared LLM, analytics, telemetry, unpinned evidence, GPU APIs, or unexpected third parties. Confirm Canvas2D, optional browser-native speech, persistent captions, truth badges, species boundaries, visible scale bars, non-overlapping labels, recognizable motor layers, and a distinct inference/Markov boundary.
Phase 7 — Deep falsification even after green gates
Form explicit null hypotheses for implementation error, data leakage, flexible fit, non-identifiability, cross-study confounding, failure of prospective prediction, alternative mechanisms, scale transfer, pedagogical analogy, and parity overclaim. For each specify falsifier, data, experimental unit, controls, sample-size rationale, stopping rule, preregistration, and analysis before revealing results.
Run parallel isolated experiment families with deterministic seeds:
- leave-one-motor/cell/study/condition/species/intervention-out and temporal forward prediction;
- identical-split comparison against constant, empirical, Markov, semi-Markov, survival, flexible spline/GAM, HMM, hierarchical Bayesian, and credible non-UNI mechanistic baselines;
- ablations of priors, slow state, policy, residual feedback, stator state, load, PMF, CheY-P proxy, lattice coupling, motor identity, and study effects;
- parameter recovery in correctly specified and misspecified synthetic worlds;
- posterior predictive checks of dwell, hazard, survival, tails, dispersion, unit variability, autocorrelation, switching asymmetry, and load dependence;
- robustness across priors, seeds, exclusions, outliers, discretization, censoring, uncertainty, tolerances, bounds, units, and structure;
- negative controls using stratified shuffles, identity shuffles, time shifts, irrelevant covariates, reversed time, broken boundaries, and conditions where the mechanism should not apply;
- frozen prospective manifests containing code/data/split/model hashes, prediction, uncertainty, score, threshold, and timestamp before reveal;
- model-disagreement mapping to find feasible, maximally discriminating future measurements;
- causal intervention designs for PMF, load, stator availability, CheY-P, temperature, viscosity, components, and ligand environment.
Map the validity domain over species, strain, motor, cell, stator count, load, PMF, temperature, viscosity, signaling, apparatus, timescale, scale, source, formulation, and parameter regime. Classify regions as supported, tentative, contradicted, unidentifiable, unobserved, or extrapolation-only. Identify where UNI beats serious alternatives, ties simpler models, fails, or requires study-specific tuning.
Rank next experiments by expected information gain, discriminating power, feasibility, independence, reproducibility, cost, and self-deception risk. For each candidate breakthrough specify hypothesis, alternatives, mathematical change, biological meaning, data, frozen prediction, sample-size rationale, accept/reject thresholds, replication, and what success would and would not establish. Prestige is not an acceptance criterion.
Modification policy
Start read-only. If a delta is found, preserve pre-fix evidence, add a failing test, prepare the smallest patch in an isolated branch, run focused and complete validation, and report rollback. Do not weaken gates or rewrite historical reports. Do not push, deploy, publish, change access, or mutate external systems without Michael's explicit authorization.
Required paste-back report
Return one Markdown report using these exact sections:
# CLAUDE / ULTRA CODE INDEPENDENT AUDIT## Executive Verdict— commit, tree state, software/science/release verdicts, strongest result, largest risk, best next experiment.## Commands Actually Run— command, environment, exit, duration, artifact.## Gate Ledger—PASS | FAIL | BLOCKED | NOT RUN | EXTERNAL VALIDATION REQUIRED.## Deltas Found— numbered deltas with severity, domain, affected claim, file/line, observed/expected behavior, evidence, reproduction, root cause, correction, failing test, risk, conclusion impact, patch status, and diff.## Surviving Mutations.## Mathematical Independent Checks.## Evidence and Provenance Findings.## Human Walkthrough Record.## Deep Falsification Results— hypothesis, alternatives, data/hash, unit, sample, split, frozen prediction, metric, uncertainty, baseline, outcome, interpretations, command, artifact/hash.## Failed and Adverse Results— this section may never be omitted.## Validity-Domain Map.## Candidate Model Breakthroughsranked by information gain.## Exact Next Actions for Codex— bounded actions with prerequisites, files, failing test, implementation, validation, evidence, rollback, and acceptance criterion.## Paste-Back Capsule— audited commit, clean/dirty state, gate counts, delta counts, top discoveries, top actions, artifact paths, and exact commands.
If no implementation delta exists, write exactly:
NO IMPLEMENTATION DELTAS FOUND IN THE EXECUTED TEST DOMAIN.
Then continue into deep falsification. “All executed gates passed” is the strongest allowed conclusion from green gates alone.
Your governing objective is:
Find the strongest world in which the model survives serious alternatives, map precisely where it fails, and let risky prospective evidence—not desire— decide whether that world can expand.
sha256 84a20fd7709ca5c7 — of the original file, so what was ingested stays checkable.
Plain — written for this website, not the source document
A prompt, not a report. It is written to be handed to an assistant that has been given access to a private repository, telling it how to audit a scientific software project from the outside. Nothing here is a finding. Everything it describes is work that would still have to be done.
The instruction at its heart is simple. Do not agree by default, and do not work toward a green dashboard. Rebuild the project in a fresh copy, run its checks, then deliberately try to break them and see whether anything notices — by relabelling a reconstruction as an observation, say, or by swapping one species for another. Redo the important mathematics by a second, independent route. Trace every claim back to the source it came from. Write down what failed, and never quietly drop a result that came out badly.
It is worth reading for anyone who would like to see what it looks like when a project asks to be taken apart rather than admired, and writes the rules for that in advance.
Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 84a20fd7709ca5c7
Clear — written for this website, not the source document
An instruction sheet for someone else to follow, and not a record of anything that happened. It is written to be pasted into an assistant that has been given access to a private repository, and it asks that assistant to audit a scientific software project as an independent outsider. It reports no results. Everything in it is work that would still have to be done.
The mission it sets is blunt. Build the repository, run it, inspect it, and try to falsify it. Do not agree by default, and do not aim at a green dashboard. Work out what the implementation and the evidence actually support. It also fences the finished product: whatever tools the reviewer uses, the released application must stay on ordinary processors and carry no hidden model calls, no tracking and no accounts. And it names the temptation it works hardest to head off, which is reading passing software tests as though they had settled the biology, or human parity, or scientific significance. Those stay separate hypotheses.
The work runs in stages. The first changes nothing: record the machine, the versions and the exact state of the copy. Next, list every claim the project makes anywhere: in documents, in the interface, in code, in captions, in reports. For each one, record where it came from, which species it concerns, what test covers it, what would show it wrong, and whether it is supported, conditional, unsupported, contradicted, not tested or external.
Then a clean rebuild, with every command, its exit status, its outputs and its checksums written down. If something needed is missing, the honest entry is that it did not run. Absence may not be converted into a pass.
Then the tests themselves go on trial. In a separate, disposable copy the reviewer breaks things on purpose. Labels reconstruction output as observed, swaps the species, leaks motors between training and holdout. Counts frames as independent replicates, removes an adverse result, computes a prediction after the observation has been revealed. Each time, the question is whether the tests catch it. Anything that survives is a gap in the tests, and the damage is never committed back.
After that, the central mathematics is rederived by a second independent route, with units, assumptions and failure conditions stated. The evidence is rechecked against its sources, so that nothing at runtime can relabel reconstruction, derived data or inference as observed. A person then walks the application by hand, on small screens, at high zoom, with a keyboard only, writing a prediction down before the evidence is revealed.
The final stage is the point of the whole page: even when every gate is green, keep trying to break it. Compare against simpler models on identical splits, remove parts to see what they were doing, run negative controls, freeze predictions before the reveal, and map where the model holds and where it fails. Rank what to measure next by expected information gain. Prestige is not an acceptance criterion.
The page closes by prescribing the exact shape of the report to be returned, including a section for failed and adverse results that it says may never be omitted.
Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is 84a20fd7709ca5c7