UNI Universal Natural Intelligence

Wiki · The Cookbook

CN-05 — DNA: the information substrate

The Cookbook · cookbook/recipes-natura/CN-05-dna.md @ 575fc93d9d31 (main) — opens the published snapshot e850f872196d

How to read this page

Three ways to read this page. Precise is the document itself, exactly as it is written in the repository. Plain and Clear were written for this website to help you meet that document — they are about it. They are not it, and they are not evidence.

The Cookbook is the method carried out step by step: 34 pages of recipes for building a developmental active-inference SIMULATION — a bounded peek at a toy world, never a person. The front matter says that word is never softened under any pressure, so it is not softened here. The recipes run from the molecular and cellular rungs up through metabolism, motor control, perception, language and metacognition, and on to rungs that are still open questions. Around them sit a set of kitchen rules, a shared pantry of engines and primitives, and a second family of recipes about nature itself — rocks, water, air, stars, DNA, ants, whales, bats, humans.

It is for the reader asking what building this would actually take. Each recipe names its ingredients, the order of work, the tests to run at that stage, and the point at which a step stops being something already carried out and becomes something proposed.

Begin with the front matter and then the kitchen rules. Those two pages fix the honest position and the fence labels that every later recipe leans on, and without them the status markers on a recipe are easy to skim past. After that the recipes can be read in any order.

The nature recipes sit slightly apart and should be read that way. They cite outside science — geology, chemistry, biology, astrophysics — and a nature citation is never a UNI gate: those chapters contain zero UNI claims and raise no rung.

What it is not: a claim that the whole ladder has been cooked. The book recommends the complete recipe and, on the same page, labels every rung by its real state — that tension is deliberate and is the thing the book is built around. Where a recipe and the claim ledger disagree, the ledger wins.

Your browser cannot switch reading levels, so the document itself is shown.

Precise — the source document

This is the document. Rendered from the repository at the commit above, with nothing rewritten for the web. A gate re-renders it on every deploy and fails the build if a single byte differs.

What you are reading. The molecule treated as what it is: a physical polymer with measured dimensions and mechanics, carrying a measurable quantity of information at a measured error rate, packed by a measured hierarchy, copied at measured speed. Geometry first, information second, and the deflation immediately after — because the single most expensive mistake available here is to read DNA as source code. Every number below carries a source you can find or is written NOT-MEASURED. Nothing here raises any UNI rung.


Geometry first

The B-form double helix, at the resolution a designer actually needs:

  • Rise: ~3.4 Å per base pair (0.34 nm). This is the number that converts sequence length into physical length, and it is the most load-bearing constant in the chapter.
  • Diameter: ~20 Å (2 nm).
  • Base pairs per turn: this is where the textbook is wrong, instructively.

Watson & Crick (1953), Nature 171:737–738, proposed the structure with 10 bp/turn — a value read off X-ray fibre diffraction of oriented, semi-crystalline DNA (Wilkins, Stokes & Wilson, Nature 171:738, and Franklin & Gosling's Photo 51 appear in the same 25 April 1953 issue; exact Franklin & Gosling pagination NOT-CONFIRMED in this pass). Fibres are not solution.

Wang (1979), PNAS 76(1):200–203, doi:10.1073/pnas.76.1.200, measured it properly in solution: run pairs of covalently closed circular DNAs differing by 1–58 bp out of ~4,350 bp on a gel and read the topoisomer ladder. Result: 10.4 ± 0.1 bp/turn, which Wang states is "significantly different from the value 10.0 base pairs per turn for the B form fiber structure."

Rhodes & Klug (1980), Nature 286(5773):573–578, PMID 7402337, "Helical periodicity of DNA determined by enzyme digestion," bound short stiff DNA to flat surfaces, digested with DNase I and read the cutting periodicity: 10.6 ± 0.1 bases.

These two do not resolve into a tidy story, and the tidy story is the trap. Rhodes & Klug do not report 10.6 as a surface-bound value. Their abstract's final sentence, verbatim: "We identify this value with the number of base pairs per turn of the DNA double helix in solution." So Wang's 10.4 and Rhodes & Klug's 10.6 are two claims about the same quantity under the same nominal condition, reached by two different assays — a topoisomer ladder against a DNase I cutting periodicity. That is a method-dependent discrepancy, and this pass does not reconcile it. Both are primary. They disagree.

The ubiquitously quoted "~10.5" sits between them. It is not nowhere: Potaman & Sinden's Table 1 — this chapter's source for the rest of the geometry — prints 10.5 as B-DNA's residues per helical turn. But why the field settled on 10.5 (an independent determination? a rounding? a midpoint?) is NOT-SOURCED in this pass. Quote 10.5 if you like; know that it is neither Wang's number nor Rhodes & Klug's, and that the two primaries it sits between have not been reconciled here.

Grooves — and a convention trap this chapter fell into. Two pairs circulate for B-DNA's major/minor groove widths: ~12 Å / ~6 Å, and the widely repeated textbook pair 22 Å / 12 Å. Both are NOT-SOURCED in this pass.

An earlier draft of this chapter attributed the 12/6 pair to Potaman & Sinden, "DNA: Alternative Conformations and Biology," Madame Curie Bioscience Database (Landes Bioscience, NCBI Bookshelf NBK6545). That citation was fabricated. NBK6545's Table 1 carries no groove-width row at all — its rows are direction of helix rotation, residues per helical turn, axial rise, pitch, base-pair tilt, rotation per residue, diameter, glycosidic bond, and sugar pucker — and its prose describes the grooves only qualitatively, as a cylinder of 20 Å diameter bearing a major and a minor groove. The 12/6 values are standard in the literature and may well be right; the receipt was not. Recorded below, not quietly deleted.

The conventions are the reason two pairs exist at all: groove width is reported either as a phosphate–phosphate distance or with the van der Waals radii of the phosphates subtracted. But watch what that does not buy. The offset usually quoted is ~5.8 Å, and it accounts for the minor pair (12 − 5.8 = 6.2 ≈ 6) while failing on the major pair: 22 − 5.8 = 16.2, not 12. The two printed pairs differ by 10 Å in the major groove. So "different convention" is not a sufficient explanation, the ~5.8 Å offset is itself NOT-SOURCED here, and the provenance of the textbook 22 Å remains genuinely unexplained. A groove width quoted without its convention is underspecified — and this chapter did not have a sourced one. Do not average them; name the convention or do not print the number. The chapter now obeys its own rule here rather than lecturing it.

The three families (same source unless noted):

A-DNA B-DNA Z-DNA
Sense right right left
bp/turn 11 10.5 (this source); cf. 10.4 soln (Wang) / 10.6 soln (Rhodes & Klug) / 10 fibre 12
Rise 2.55 Å 3.4 Å 3.7 Å
Diameter 23 Å (9 Å axial hole) 20 Å 18 Å
Sugar pucker C3′-endo C2′-endo alternating

Z-DNA is not a curiosity invented to fill a table — it was solved crystallographically at atomic resolution: Wang et al. (1979), Nature 282:680, "Molecular structure of a left-handed double helical DNA fragment at atomic resolution." The molecule has more than one stable conformation, and which one it occupies depends on sequence and environment. A design that treats DNA as one fixed geometry has already lost information.

It is a polymer, so it has mechanics

The number that turns DNA from a diagram into an object: persistence length ℓ_p ≈ 50 nm ≈ 150 bp in ~0.1 M aqueous NaCl at ~20–25 °C, a consensus value reproduced by magnetic tweezers, optical tweezers, and AFM (see the review by Peters & Maher, "DNA curvature and flexibility in vitro and in vivo," Q Rev Biophys 43(1):23–63, PMID 20478077).

The temperature belongs in that scope, and this chapter previously omitted it. Geggier, Kotlyar & Vologodskii (2011), NAR 39(4):1419–1426, doi:10.1093/nar/gkq932, PMID 20952402, "Temperature dependence of DNA persistence length," measured ℓ_p falling as it warms: 53.2 nm at 5 °C → 42.5 nm at 42 °C. Their conclusion, verbatim: DNA persistence length "strongly depends on temperature and accounting for this dependence is important in quantitative comparison between experimental results obtained at different temperatures." Two consequences, and both cut against citing this paper as support for the 50 nm figure. First, they worked in TBE with 10 mM MgCl₂ — not ~0.1 M NaCl — so it is a different ionic condition, a condition on the consensus rather than a replication of it. Second, their values at and above body temperature sit at or below the 45–55 nm band this chapter prints as its own falsifier. By the chapter's own rule that a rate without a temperature is underspecified, a persistence length without one is too.

Read it as an engineer. Below ~150 bp, DNA is effectively a stiff rod — you cannot bend it appreciably with thermal energy. Above it, DNA is a flexible coil and behaves like a worm-like chain. This single constant explains why a nucleosome is a problem: wrapping 147 bp around a histone octamer means bending DNA through ~1.65 superhelical turns at its persistence length, which is not free. Something has to pay, and the histones pay it with binding energy.

Fence the constant honestly: sub-100 bp DNA is more bendable than the worm-like chain predicts — Vafabakhsh & Ha, Science (PMC3565842), report extreme bendability by single-molecule cyclisation. The 50 nm value is a large-scale property, not a licence to extrapolate to short loops.

The information content, then the deflation

Four bases, equiprobable: log₂(4) = 2 bits/bp. That is the ceiling, and it assumes no correlations — real genomes have plenty, so 2 bits/bp is an upper bound, not a content measurement.

Use the honest genome size. Not GRCh38 ("~3.1 Gbp" with gaps), but T2T-CHM13: Nurk et al. (2022), Science 376:44–53, doi:10.1126/science.abj6987 — 3,054,815,472 bp of gapless nuclear DNA (chromosomes 1–22 and X) plus a 16,569 bp mitochondrial genome. Now compute:

Haploid : 3.055e9 bp x 2 bits = 6.11e9 bits = 764 MB  (728 MiB)
Diploid : ~6.11e9 bp x 2 bits = 1.22e10 bits = 1.53 GB (1.42 GiB)

A human genome fits on a DVD. Hold that number, then deflate it immediately, because it is not a design specification.

Coding DNA is 1–2% of the genome (Piovesan et al. 2019, BMC Res Notes, "Human protein-coding genes and gene feature statistics in 2019," doi:10.1186/s13104-019-4343-8; the exome is ~1.5%). ~20,000 protein-coding genes. Whatever the other 98% is doing, it is not spelling out a body plan.

And then the C-value paradox finishes the job. Genome size does not track organismal complexity — not weakly, not noisily, not at all (conversion throughout: 1 pg = 0.978 Gbp, Doležel et al. 2003, Cytometry A, doi:10.1002/cyto.a.10013):

Organism Genome (1C) vs human
Human (T2T-CHM13) 3.055 Gbp
Onion, Allium cepa 16.75 pg ≈ 16.4 Gbp; ≥95% repetitive 5.4×
South American lungfish, Lepidosiren paradoxa 91 Gbp, ~90% repeat content — largest animal genome sequenced ~30×
Paris japonica 152.23 pg ≈ 149 Gbp ~49×
Fork fern, Tmesipteris oblanceolata 160.45 Gbp — current eukaryotic record 52×

Sources: Fernández et al. (2024), iScience, "A 160 Gbp fork fern genome shatters size record for eukaryotes," doi:10.1016/j.isci.2024.109889; Schartl et al. (2024), Nature 634:96–103, doi:10.1038/s41586-024-07830-1 (lungfish); Pellicer, Fay & Leitch (2010), Bot J Linn Soc 164(1):10, doi:10.1111/j.1095-8339.2010.01072.x (Paris japonica).

A fern carries 52× your information budget. If genome size measured design content, that fact would be a catastrophe. It is not a catastrophe; it is a refutation of the premise.

The error budget, in layers — the chapter's best engineering lesson

Routinely conflated, so do it carefully. The famous ~10⁻¹⁰ is not the polymerase's error rate. It is what comes out of a stack.

Schaaper (1993), J Biol Chem 268(32):23762–23765, "Base selection, proofreading, and mismatch repair during DNA replication in Escherichia coli" (PMID 8226906), measured the three layers in E. coli by sequencing 866 lacI mutations across strains with the error-correction pathways disabled. What he reports are fold-discriminations, not power-of-ten error rates — verbatim, base selection discriminates against errors by 200,000–2,000,000-fold, proofreading by 40–200-fold, and mismatch repair by 20–400-fold, "each depending on the type of error."

The familiar decade decomposition is a different paper, and this chapter previously misattributed it to Schaaper. It is Fijalkowska, Schaaper & Jonczyk (2012), FEMS Microbiol Rev 36(6):1105–1121, doi:10.1111/j.1574-6976.2012.00338.x, PMID 22404288 — a review, citing Schaaper — which reads, verbatim: "Roughly, the contribution of each of these components to the error rate can be estimated at 10⁻⁵ (insertion), 10⁻² (proofreading), and 10⁻³ (mismatch repair), accounting for the 10⁻¹⁰ overall rate, although each contribution is highly dependent on the precise type of error under question." Its abstract puts the overall rate at 10⁻⁹ to 10⁻¹¹ errors per base pair.

Keep the two apart, because the decades do not reproduce the measurement. Proofreading at 40–200-fold is 10⁻¹·⁶–10⁻²·³; MMR at 20–400-fold is 10⁻¹·³–10⁻²·⁶. The round decades are the review's convenience — reasonable, explicitly hedged by its own "roughly," and not Schaaper's numbers. Kunkel's reviews (Kunkel 2004, J Biol Chem 279(17):16895–16898; Kunkel & Bebenek 2000, Annu Rev Biochem 69:497–529) give in vitro polymerase insertion fidelity at ~10⁻⁴–10⁻⁵, with 3′→5′ exonucleolytic proofreading and then MMR stacked on top.

Put the diploid genome through it (6.11 × 10⁹ bp, one replication):

Layer Error rate/bp Errors per replication
Polymerase base selection alone ~10⁻⁵ (in vivo estimate) — cf. ~10⁻⁴ in vitro ~61,000cf. ~610,000 at the in-vitro rung
+ 3′→5′ exonucleolytic proofreading (× ~10⁻²) ~10⁻⁷ ~610
+ mismatch repair (× ~10⁻³) ~10⁻¹⁰ ~0.6

Read the ladder's own arithmetic before you trust it. This table uses Fijalkowska et al.'s decomposition, which has the merit of actually composing: 10⁻⁵ × 10⁻² × 10⁻³ = 10⁻¹⁰ exactly. An earlier draft of this table started at the in vitro 10⁻⁴ rung and still printed 10⁻⁷ after proofreading — which silently requires a 10³ gain, contradicting the same row's own stated gain of 10¹–10², contradicting the text above it, and contradicting Schaaper's measured 40–200-fold. That rung was arithmetically broken and is corrected here. Nor does the corrected ladder escape the deeper problem: composing Schaaper's measured factors onto a 10⁻⁵ rung spans ~10⁻⁷·⁹ to ~10⁻⁹·⁹, reaching 10⁻¹⁰ only at its most generous edge. The decades are round because they were chosen to land on 10⁻¹⁰ — not because three measurements happened to multiply out that way. The ladder is MODELED at every rung, and its rungs are the review's estimate, not any single paper's measurement.

The principle survives that, and it generalises far beyond biology: nature did not build one perfect component. It stacked three cheap imperfect ones, each individually unimpressive, and multiplied their independent failure probabilities. A ~10⁻⁵ part, a ~10⁻² filter, and a ~10⁻³ filter compose to 10⁻¹⁰. No single element in the chain is remarkable. The chain is. If you set out to build a 10⁻¹⁰ component you will fail; if you set out to build three cheap layers whose failures are uncorrelated, you will succeed. The independence is the whole trick — correlated failures collapse the product back toward the worst layer. Note what that argument does and does not need: it needs the layers to be cheap, stacked, and independent. It does not need any particular rung, which is fortunate, because the rungs are the softest numbers in this chapter.

Now fence the first rung, because it is contested. St Charles et al. (2015), DNA Repair (Amst) 31:41–51, doi:10.1016/j.dnarep.2015.04.006, dissected the layers in vivo in yeast and report, verbatim: in the absence of proofreading and MMR, Pol ε and Pol δ synthesise DNA in vivo "with apparent base selectivity that is more than 100 times higher than measured in vitro." They further find proofreading is strand-asymmetric (Pol ε on the leading strand vs Pol δ on the lagging strand contribute differently), and that "on average, proofreading contributes more to replication fidelity than does MMR" — while the per-mismatch split "vary[s] from nearly all proofreading of some mismatches to mostly MMR of other mismatches."

So the tidy 10⁻⁵ → 10⁻⁷ → 10⁻¹⁰ ladder is a teaching model assembled from measurements taken in different systems against different denominators. Three positions now sit on the first rung alone: ~10⁻⁴ in vitro (Kunkel); ~10⁻⁵ as a review's in vivo estimate (Fijalkowska et al. 2012); and, measured in vivo in yeast, base selectivity >100× better than in vitro (St Charles et al. 2015) — which would push the rung below 10⁻⁶. This pass did not reconcile them and does not print a reconciliation. The stacking principle survives the discrepancy intact; the specific rungs do not travel without their assay.

One more denominator trap: the human germline rate of 1.20 × 10⁻⁸ per nucleotide per generation (Kong et al. 2012, Nature 488(7412):471–475, doi:10.1038/nature11396, at mean paternal age 29.7, with paternal mutations doubling every ~16.5 years) is per generation, not per replication. A generation contains many germline divisions. Multiplying it by 6.11 Gbp gives ~73 de novo sites per generation as a MODELED figure — but do not set it beside 10⁻¹⁰ as if they measured the same thing.

Packing: 2 metres into 6 micrometres

Do not cite the 2 m; derive it, from the rise:

3.055e9 bp x 0.34 nm/bp = 1.04 m  (haploid)
                        ~ 2.08 m  (diploid)
Nucleus, d = 6 um:  V = (4/3)pi(3 um)^3 = 113 um^3
Linear packing ratio: 2.08 m / 6e-6 m ~ 3.5e5

But run the volume, which almost nobody does. DNA as a cylinder of radius 1 nm:

V_DNA = pi (1e-3 um)^2 x 2.08e6 um = 6.5 um^3  ->  6.5 / 113 = ~5.8% of nuclear volume

The nucleus is 94% not-DNA. Packing is not a space problem. It is a topology and access problem: keep 2 m of a 50-nm-persistence-length polymer untangled, unbroken, and selectively readable. Volume was never the constraint. That reframing is the point.

The hierarchy, with its honesty applied where it is owed:

  • Nucleosome — solid. Luger et al. (1997), Nature 389(6648):251–260, doi:10.1038/38444, solved the core particle at 2.8 Å with 146 bp of α-satellite DNA wrapped 1.65 left-handed superhelical turns around a histone octamer. The canonical 147 bp comes from the higher-resolution follow-up: Richmond & Davey (2003), Nature 423:145, doi:10.1038/nature01595, at 1.9 Å. Both numbers are real; they are different crystals. (Richmond & Davey also found nucleosomal DNA carries "twice the curvature necessary" for the superhelical path — the bending is not gentle.)
  • The 30 nm fibre — CONTESTED in vivo, and this matters. It is in every textbook. Maeshima, Hihara & Eltsov (2010), Curr Opin Cell Biol, doi:10.1016/j.ceb.2010.03.001 (PMID 20346642), asked directly — "Chromatin structure: does the 30-nm fibre exist in vivo?" — and report that cryo-EM of vitrified human mitotic cells, imaged close to native state, found no 30-nm fibres. Ou et al. (2017), Science 357(6349):eaag0025, doi:10.1126/science.aag0025, using ChromEMT electron tomography, found chromatin to be "a disordered 5- to 24-nanometer-diameter curvilinear chain" packed at varying 3D concentration — again no 30-nm fibre. (ChromEMT requires fixation, dehydration, heavy-metal staining and plastic embedding; cryo-ET does not. The two methods have different artifacts and agree anyway, which is why this is strong.) The 30 nm fibre is well established in vitro. Its existence as an in vivo structural level is OBSERVED-CONTESTED and should not be drawn in a diagram without the fence.
  • Loops / TADs / chromosome — the levels above. NOT-SOURCED in this pass.

Replication: the arithmetic that makes a huge genome copyable

Rates first (Milo & Phillips, Cell Biology by the Numbers, "How long does it take cells to copy their genomes?"):

  • E. coli fork: ~600 bp/s in vivo average (BNID 109251); classic figure ~1,000 nt/s.
  • Eukaryotic fork: 4–40 bp/s ≈ 1 kb/min (BNID 104930, 104935, 104936, 104937). ~25–150× slower.

E. coli, 4.6 Mbp, one origin, bidirectional: 4.6e6 / (2 x 600) = 3,833 s = 64 min (at 1,000 bp/s: 38 min). But E. coli divides in ~20 min (BNID 103514) — faster than it can copy itself. The trick is overlapping replication cycles: fire new origins before the last round finishes, running >6 origins and >10 forks at once (BNID 102356). The cell is replicating a genome it has not finished replicating.

Humans cannot use that trick, and the fork is far slower. Do the arithmetic. Diploid 6.11 Gbp, one origin, two forks, v = 20 bp/s:

6.11e9 / (2 x 20) = 1.53e8 s = 4.8 YEARS

Measured S phase is ~10 hours (BNID 103742, 103741, 102204). Solve for the origins required:

N = 6.11e9 / (2 x 20 bp/s x 3.6e4 s) ~ 4,200 origins
   (v = 4 bp/s -> ~21,000;  v = 40 bp/s -> ~2,100)

Measured human origin count: ~1,000 to 100,000 (BNID 107654, 109283). The prediction lands inside the measurement. 4.8 years → 10 hours, a factor of ~4,200, bought entirely by parallelism at bidirectional origins — not by a faster polymerase. (Pushed to the limit: D. melanogaster embryos copy ~120 Mbp every 8 minutes, BNID 101971.) Nature's answer to "this component is too slow" was never "build a faster component."

The code: 64 → 20, and why the redundancy is not noise

4³ = 64 codons; 61 sense codons for 20 amino acids plus 3 stops. Average 3.05 codons per amino acid. Information-theoretically the code throws away log₂(64) − log₂(21) = 6 − 4.39 = 1.61 bits per codon.

That discarded 1.61 bits is not waste — it is error tolerance, and it is measurably non-random. Freeland & Hurst (1998), J Mol Evol 47(3):238–248, doi:10.1007/PL00006381, PMID 9732450, scored the canonical code against randomly generated alternatives for how well it minimises the damage of point mutation and mistranslation (errors land on synonymous codons or on chemically similar amino acids, measured by polar requirement). Verbatim from the abstract: "if we employ weightings to allow for biases in translation, then only 1 in every million random alternative codes generated is more efficient than the natural code."

The unweighted figure is a different paper, seven years earlier — and it is not quite the number everyone quotes. Haig & Hurst (1991), J Mol Evol 33(5):412–417, doi:10.1007/BF02103132, PMID 1960738, "A quantitative measure of error minimization in the genetic code," report — verbatim — that "single-base changes in the natural code had a smaller average effect on polar requirement than all but 0.02% of random codes." 0.02% is ≈1 in 5,000, not 1 in 10,000. The familiar 10⁻⁴ headline is a secondary restatement: Koonin & Novozhilov (2009), IUBMB Life 61(2):99–111, PMID 19117371, write that "the probability of a random code to be fitter than the standard one is P₁ ≈ 10⁻⁴" — a ~2× loosening of the primary it compresses. Freeland & Hurst 1998's abstract contains no 1-in-10,000 figure at all; an earlier draft of this chapter sourced both endpoints to that one paper, which misattributed Haig & Hurst's result.

So the famous pair is two papers, two conditions, and one rounding — not one paper's two settings. The "one in a million" arrives only once transition/transversion bias and mistranslation bias are weighted in. Carry the conditions or the headline is not the result — and carry the citation, or the headline is not even the right author's.

The honest later critique, which is essential. Novozhilov, Wolf & Koonin (2007), Biol Direct, "Evolution of the genetic code: partial optimization of a random code for robustness to translation error in a rugged fitness landscape" (PMID 17956616), find the standard code is "the result of partial optimization of a random code" — highly robust, "but there is a huge number of more robust codes," and it could evolve from a random code via a short series of codon reassignments. So: the code is a local optimum on a rugged landscape, not the global optimum, and not a miracle. "One in a million" and "not the best available" are both true. That pair is the whole discipline of this chapter in one line.

The central dogma, and its real exceptions

DNA → RNA → protein, with information not flowing back out of protein. The exceptions are real, sourced, and not merely decorative:

  • Reverse transcription. Temin & Mizutani (1970), Nature 226:1211–1213, doi:10.1038/2261211a0, and Baltimore (1970), Nature 226:1209–1211 — back-to-back in the same 27 June 1970 issue. RNA → DNA. The arrow reverses.
  • Prions. Prusiner (1982), Science 216:136–144, doi:10.1126/science.6801762, PMID 6801762: "a small proteinaceous infectious particle which is resistant to inactivation by most procedures that modify nucleic acids." Heritable conformational information carried with no nucleic acid. The substrate itself is optional.

DNA is a hereditary substrate. It is not the hereditary substrate.

The active-inference reading — HYPOTHESIZED, a lens only

Fenced explicitly. This is a framing, not a measurement, and no observation below is offered as evidence for it.

One may read the genome as a prior over phenotypes, fitted by ancestral surprise: lineages whose priors mispredicted their niche were deleted, so the surviving distribution encodes something about the statistics of the environments that did the deleting. The layered error budget then reads as precision control — the accuracy of transmission set to a value, not maximised, since a 10⁻¹⁰ genome and a 10⁻³–10⁻⁴ ribosome coexist in one cell by different budgets.

This is a lens. It generates questions (what would a measured prior look like? what units?). It measures nothing, predicts no number in the table above, and could be deleted from this chapter without changing a single value. Treat it accordingly. A UNI design must not cite it as support for anything.

The numbers

Symbol Value Units Scope Class Source Falsifier
rise ~3.4 (0.34) Å (nm) /bp B-DNA OBSERVED-REPLICATED Potaman & Sinden, NCBI Bookshelf NBK6545; Watson & Crick 1953, Nature 171:737–738 Structural measurement outside 3.3–3.5 Å under B-form conditions
bp/turn (fibre) 10.0 bp B-form fibre diffraction OBSERVED-REPLICATED Watson & Crick 1953; as compared in Wang 1979 Re-analysis of fibre data giving ≠10
bp/turn (solution) 10.4 ± 0.1 bp B-DNA free in solution, physiological; topoisomer gel OBSERVED-REPLICATED Wang 1979, PNAS 76(1):200–203, doi:10.1073/pnas.76.1.200 Repeat topoisomer ladder; value outside 10.3–10.5
bp/turn (Rhodes & Klug) 10.6 ± 0.1 bases DNase I cutting periodicity on DNA immobilised on three surfaces; the authors identify the value with the repeat in solution OBSERVED-REPLICATED Rhodes & Klug 1980, Nature 286(5773):573–578, PMID 7402337 Repeat digestion periodicity outside 10.5–10.7
bp/turn 10.4 vs 10.6 unreconciled — two primaries, same nominal condition (solution), two assays bp topoisomer ladder (Wang) vs DNase I periodicity (Rhodes & Klug) OBSERVED-CONTESTED Wang 1979; Rhodes & Klug 1980 An assay reconciling the two, or a re-measurement collapsing the gap
bp/turn "10.5" 10.5 bp quoted everywhere; printed by Potaman & Sinden Table 1 as B-DNA residues/turn; sits between two unreconciled primaries NOT-SOURCED in this pass (value printed in source; provenance not traced) Potaman & Sinden, NBK6545 Table 1 Establish whether 10.5 is an independent determination, a rounding, or a midpoint of 10.4 and 10.6
diameter 23 (A) / ~20 (B) / 18 (Z) Å A-, B-, Z-DNA OBSERVED-REPLICATED Potaman & Sinden, NBK6545 Table 1 Structural measurement outside range
groove widths (12/6) major ~12, minor ~6 Å B-DNA, convention not sourced here NOT-SOURCED in this pass not confirmed here — NBK6545 Table 1 has no groove-width row; the prior attribution to it was fabricated (recorded below) Fetch a primary printing these values and naming the measurement convention
groove widths (22/12) 22 / 12 Å widely repeated textbook pair — different convention NOT-SOURCED in this pass not confirmed here Fetch a primary source and name the measurement convention
groove convention offset ~5.8 Å claimed offset between P–P and vdW-corrected conventions NOT-SOURCED in this pass not confirmed here Fetch a primary defining groove width as smallest P–P separation minus the phosphate diameter. Note it accounts for 12 → 6.2 ≈ 6 but fails on 22 → 16.2 ≠ 12 (a 10 Å gap), so it does not reconcile the two pairs
A-DNA 11 bp/turn; rise 2.55 Å; C3′-endo; 9 Å axial hole; tilt ~20° A-form OBSERVED-REPLICATED Potaman & Sinden, NBK6545 Structural measurement outside range
Z-DNA left-handed; 12 bp/turn; rise 3.7 Å; 18 Å; 30°/bp Z-form, atomic-resolution crystal OBSERVED-REPLICATED Wang et al. 1979, Nature 282:680; Potaman & Sinden Re-refinement contradicting handedness or repeat
ℓ_p ~50 (~150) nm (bp) dsDNA, ~0.1 M NaCl, ~20–25 °C; tweezers + AFM consensus OBSERVED-REPLICATED Peters & Maher 2010, "DNA curvature and flexibility in vitro and in vivo," Q Rev Biophys 43(1):23–63, PMID 20478077 Independent single-molecule measurement outside 45–55 nm at the stated temperature and ionic strength
ℓ_p vs temperature 53.2 nm (5 °C) → 42.5 nm (42 °C) nm TBE + 10 mM MgCl₂ — a different ionic condition from the row above; j-factor + linking-number methods. A condition on the 50 nm consensus, not support for it: at ≥37 °C it sits at or below that row's own falsifier band OBSERVED-REPLICATED Geggier, Kotlyar & Vologodskii 2011, NAR 39(4):1419–1426, doi:10.1093/nar/gkq932, PMID 20952402 A measurement showing ℓ_p temperature-independent across 5–42 °C
ℓ_p sub-100 bp short DNA more bendable than WLC predicts <100 bp cyclisation OBSERVED-CONTESTED Vafabakhsh & Ha, Science, PMC3565842 A cyclisation method restoring WLC agreement below 100 bp
I/bp 2 bits/bp log₂(4); upper bound, assumes no correlation MODELED Computed here Arithmetic error (the bound itself is definitional)
G_human 3,054,815,472 (+16,569 mtDNA) bp T2T-CHM13, gapless, chr1–22 + X OBSERVED-REPLICATED Nurk et al. 2022, Science 376:44–53, doi:10.1126/science.abj6987 Independent T2T assembly differing >0.1%
I_genome 6.11e9 bits ≈ 764 MB (haploid); ~1.53 GB (diploid) bits/bytes at 2 bits/bp, T2T-CHM13 MODELED Computed here Arithmetic error
f_coding 1–2 (exome ~1.5) % of genome human OBSERVED-REPLICATED Piovesan et al. 2019, BMC Res Notes, doi:10.1186/s13104-019-4343-8 Annotation revision moving coding fraction outside 1–2%
pg→bp 1 pg = 0.978e9 bp flow-cytometry conversion OBSERVED-REPLICATED Doležel et al. 2003, Cytometry A, doi:10.1002/cyto.a.10013 Re-derivation of nucleotide-pair molecular weight
G_onion 16.75 pg ≈ 16.4 (≥95% repetitive) Gbp Allium cepa 1C OBSERVED-REPLICATED onion assembly literature (PMC8496297; PMC11865573) + Doležel conversion Flow-cytometry re-measurement outside 15–18 Gbp
G_lungfish 91 (~90% repeat) Gbp Lepidosiren paradoxa — largest sequenced animal genome OBSERVED-REPLICATED Schartl et al. 2024, Nature 634:96–103, doi:10.1038/s41586-024-07830-1 Independent assembly differing >10%
G_P.aethiopicus ~130 Gbp marbled lungfish — estimate, NOT sequenced OBSERVED-CONTESTED / NOT-SEQUENCED flagged as unsequenced in Schartl et al. 2024 coverage Sequence it
G_fern 160.45 Gbp/1C Tmesipteris oblanceolata — current eukaryotic record OBSERVED-REPLICATED Fernández et al. 2024, iScience, doi:10.1016/j.isci.2024.109889 Independent flow cytometry differing >10%
G_Paris 152.23 pg ≈ 149 Gbp Paris japonica 1C OBSERVED-REPLICATED Pellicer et al. 2010, Bot J Linn Soc 164(1):10, doi:10.1111/j.1095-8339.2010.01072.x Re-measurement outside range
G_Polychaos 670 pg (~655 Gbp) Polychaos dubiumDO NOT USE INADMISSIBLE BNID 104470, flagged "dubious report / outdated value" Re-measure single nuclei with modern methods
ε_selection ~10⁻⁴–10⁻⁵ (in vitro); ~10⁻⁵ (in vivo E. coli, a review's round estimate) per bp polymerase base selection alone OBSERVED-CONTESTED (assay-dependent) Kunkel 2004, JBC 279(17):16895–8 (in vitro); Fijalkowska, Schaaper & Jonczyk 2012, FEMS Microbiol Rev 36(6):1105–1121, PMID 22404288 (the 10⁻⁵ estimate). Schaaper 1993 measures a 200,000–2,000,000-fold discrimination, not a rate See in-vivo row below; an assay reconciling both
Schaaper's measured factors base selection 200,000–2,000,000×; proofreading 40–200×; MMR 20–400× fold discrimination E. coli, 866 sequenced lacI mutations in correction-deficient strains OBSERVED-REPLICATED Schaaper 1993, JBC 268(32):23762–23765, PMID 8226906 A re-dissection outside these fold ranges
ε_selection in vivo vs in vitro in vivo base selectivity >100× higher than in vitro yeast Pol ε / Pol δ, proofreading- and MMR-deficient background OBSERVED-CONTESTED St Charles et al. 2015, DNA Repair 31:41–51, doi:10.1016/j.dnarep.2015.04.006 An in-vitro assay reproducing the in-vivo selectivity
ε_proof ~10⁻⁷ as a ladder rung (= 10⁻⁵ × the review's ~10⁻² factor) per bp + 3′→5′ exonucleolytic proofreading MODELED — a review's round decade; not reproduced by the measured factor rung from Fijalkowska et al. 2012; measured factor = 40–200× (10⁻¹·⁶–10⁻²·³), Schaaper 1993, JBC 268(32):23762–23765 An in vivo proofreading gain measured outside 40–200×. Prior draft printed ~10⁻⁷ off a 10⁻⁴ rung with a stated gain of 10¹–10² — arithmetically impossible (needs 10³); corrected
ε_MMR ~10⁻¹⁰ overall as the ladder rung (× the review's ~10⁻³ factor); measured overall band 10⁻⁹–10⁻¹¹ per bp + mismatch repair; overall MODELED (the ladder rung); OBSERVED-REPLICATED (the overall band) Fijalkowska et al. 2012 (abstract: "as low as 10⁻⁹ to 10⁻¹¹ errors per base pair"); measured MMR factor = 20–400× (10⁻¹·³–10⁻²·⁶), Schaaper 1993 Mutation-accumulation whole-genome rate outside 10⁻⁹–10⁻¹¹
ladder vs measured factors Schaaper's measured factors on a 10⁻⁵ rung span ~10⁻⁷·⁹ to ~10⁻⁹·⁹ — reaching 10⁻¹⁰ only at the most generous edge per bp the decades do not reproduce from the measurement MODELED Computed here from Schaaper 1993 + Fijalkowska et al. 2012 Arithmetic error; or a dissection whose measured factors compose to 10⁻¹⁰
proofreading asymmetry strand-asymmetric (Pol ε leading vs Pol δ lagging); proofreading > MMR on average, but varies per mismatch yeast, in vivo OBSERVED-REPLICATED St Charles et al. 2015 A dissection showing strand symmetry
errors/replication ~6.1e4 → ~610 → ~0.6 errors per diploid genome copy 6.11 Gbp at 10⁻⁵ / 10⁻⁷ / 10⁻¹⁰ (in-vivo ladder); ~6.1e5 at the in-vitro 10⁻⁴ rung MODELED Computed here Arithmetic error, or a refuted ε row
μ_germline 1.20e-8 per nt per generation (mean paternal age 29.7; +~2 mutations/yr; paternal doubling ~16.5 yr) human trios — different denominator from ε OBSERVED-REPLICATED Kong et al. 2012, Nature 488(7412):471–475, doi:10.1038/nature11396 Independent trio study outside range at matched paternal age
n_denovo ~73 sites/generation 1.2e-8 x 6.11e9 MODELED Computed here Arithmetic error; direct counts use a callable-fraction denominator
L_DNA 1.04 (2.08) m, haploid (diploid) 3.055e9 bp x 0.34 nm MODELED Computed here from rise + G_human Arithmetic error, or rise refuted
V_nucleus ~113 µm³ sphere, d = 6 µm MODELED Computed here (geometry) Geometric error; nuclei are not spheres
f_DNA,vol ~5.8 % of nuclear volume DNA as r = 1 nm cylinder, 2.08 m, in 113 µm³ MODELED Computed here Arithmetic error, or a refuted radius
packing ratio ~3.5e5 dimensionless (linear) 2.08 m / 6 µm MODELED Computed here Arithmetic error
nucleosome 146 bp @ 2.8 Å; 147 bp @ 1.9 Å; 1.65 superhelical turns; histone octamer bp crystal structures — different crystals, both real OBSERVED-REPLICATED Luger et al. 1997, Nature 389:251–260, doi:10.1038/38444; Richmond & Davey 2003, Nature 423:145, doi:10.1038/nature01595 A re-refinement changing the wrap length
30 nm fibre in vivo contested — not observed in cryo-EM of vitrified cells nor in ChromEMT nm in vivo interphase/mitotic chromatin OBSERVED-CONTESTED Maeshima et al. 2010, Curr Opin Cell Biol, PMID 20346642; Ou et al. 2017, Science 357:eaag0025, doi:10.1126/science.aag0025 A near-native in-vivo imaging method resolving regular 30-nm fibres
chromatin chain in vivo "disordered 5- to 24-nanometer-diameter curvilinear chain" nm ChromEMT, interphase + mitosis OBSERVED-REPLICATED Ou et al. 2017 Independent tomography contradicting the diameter distribution
loops / TADs / chromosome levels above the chain NOT-SOURCED in this pass Fetch Hi-C / loop-extrusion primaries
v_fork,ec ~600 (classic ~1,000) bp/s E. coli, in vivo average OBSERVED-REPLICATED BNID 109251; Milo & Phillips In vivo measurement outside 400–1,000 bp/s
v_fork,euk 4–40 (~1 kb/min) bp/s eukaryotic replisome OBSERVED-REPLICATED BNID 104930, 104935, 104936, 104937 Outside range
t_ec 64 min (at 600 bp/s); 38 min (at 1,000) min 4.6 Mbp, 1 origin, 2 forks MODELED Computed here Arithmetic error
E. coli multi-fork >6 origins, >10 forks; doubling ~20 min < copy time fast growth — overlapping cycles OBSERVED-REPLICATED BNID 102356; BNID 103514 A fast-growing strain with a single round per division
t_human,1origin 4.8 years 6.11 Gbp, 1 origin, 2 forks, 20 bp/s MODELED Computed here Arithmetic error
T_S ~10 hours human S phase OBSERVED-REPLICATED BNID 103742, 103741, 102204 Cell type outside range
N_origins (required) ~4,200 (2,100–21,000 over v = 40–4 bp/s) origins solved from G/(2vT) MODELED Computed here Arithmetic error
N_origins (measured) 1,000–100,000 (Drosophila ~10,000) origins human OBSERVED-REPLICATED BNID 107654, 109283 A measurement excluding the predicted band
t_Dmel ~8 min per ~120 Mbp genome D. melanogaster embryo OBSERVED-REPLICATED BNID 101971 Outside range
codons 64 → 61 sense + 3 stop → 20 aa; 3.05 codons/aa canonical code OBSERVED-REPLICATED standard; Freeland & Hurst 1998 A canonical-code recount
bits discarded 1.61 bits/codon log₂(64) − log₂(21) MODELED Computed here Arithmetic error
code optimality (unweighted) natural code beats all but 0.02% of random codes on polar requirement (≈1 in 5,000) polar-requirement metric; all single-base errors equiprobable OBSERVED-REPLICATED Haig & Hurst 1991, J Mol Evol 33(5):412–417, doi:10.1007/BF02103132, PMID 1960738 Re-run the unweighted simulation; a figure outside 0.02%
code optimality — the "1 in 10⁴" headline P₁ ≈ 10⁻⁴ a secondary restatement of Haig & Hurst, ~2× looser than their 0.02%; not a figure in Freeland & Hurst 1998, to which this chapter previously misattributed it OBSERVED-REPLICATED (as a restatement, not a primary) Koonin & Novozhilov 2009, IUBMB Life 61(2):99–111, PMID 19117371 Locate a primary reporting 10⁻⁴ directly, or retire the headline in favour of 0.02%
code optimality (weighted) 1 in 10⁶ random codes beat it + transition/transversion bias + mistranslation bias weighted in OBSERVED-REPLICATED Freeland & Hurst 1998, J Mol Evol 47(3):238–248, doi:10.1007/PL00006381, PMID 9732450 Re-run the simulation with the stated weightings
code = partial optimum "huge number of more robust codes" exist; standard code = partial optimisation of a random code on a rugged landscape OBSERVED-REPLICATED Novozhilov, Wolf & Koonin 2007, Biol Direct, PMID 17956616 A search failing to find more robust codes
reverse transcription RNA → DNA Rous sarcoma virus / RNA tumour viruses OBSERVED-REPLICATED Temin & Mizutani 1970, Nature 226:1211–1213, doi:10.1038/2261211a0; Baltimore 1970, Nature 226:1209–1211
prion heritable conformational information, no nucleic acid required scrapie agent OBSERVED-REPLICATED Prusiner 1982, Science 216:136–144, doi:10.1126/science.6801762 A nucleic acid found necessary for infectivity
genome-as-prior active-inference reading HYPOTHESIZED this chapter, as a lens Specify a measurable prior with units, then measure it

Falsifier (operable)

The chapter's central structural claim — that DNA's extraordinary system-level properties are bought by composing cheap, individually unimpressive layers whose failures are uncorrelated, not by building excellent components — is refuted by exhibiting either:

  1. a single replication component that achieves ≤10⁻⁹ errors/bp on its own, with proofreading and mismatch repair genetically ablated, at physiological rate; or
  2. a eukaryotic cell copying a ≥1 Gbp genome within one measured S phase from ≤10 origins, i.e. buying the throughput with fork speed rather than parallelism.

Either result moves the chapter. Neither Thiomargarita-style redescription nor a faster polymerase variant qualifies: the claim is about where the performance comes from, and it dies only if performance is shown to come from one place.

Secondary falsifiers are row-local: any number in the table found outside its stated scope under its stated assay moves that row and only that row. A refuted row does not refute the chapter; a refuted layering principle does.

Recorded INADMISSIBLE / NEGATIVE (first-class, inline)

  • G_Polychaos = 670 pg (~655 Gbp) as "the largest genome"INADMISSIBLE. Still the top hit in popular sources; still wrong to cite. Receipt of failure: BNID 104470 carries it flagged "dubious report" and "outdated value"; the measurement used 1960s methods assaying whole cells rather than isolated nuclei, and has never been repeated with modern methods. Excluding it, the records are Tmesipteris oblanceolata (160.45 Gbp) and Lepidosiren paradoxa (91 Gbp, sequenced). Recorded, not mocked — it was an honest 1960s measurement, and the defect is in re-citing it in 2026, not in having made it.
  • "DNA is a program / source code / blueprint"INADMISSIBLE as stated. No falsifier accompanies it; it is a metaphor doing argumentative work. The receipt is in this chapter's own numbers: 1–2% coding; a 52× genome in a fork fern; ≥95% repeat content in an onion. Keller, The Century of the Gene (Harvard UP, 2000), argues the cell may as well be read as the program and DNA as (part of) the data; Peluffo (2015), Genetics 200(3):685–696, doi:10.1534/genetics.115.178418, traces the metaphor's genesis and its failures — it omits temporality, mechanical forces in development, symbiosis, and environment. A GPT that thinks DNA is source code will design wrong: it will look for a specification where there is a prior, expect compile-time determinism where there is context-dependent chemistry, and treat 98% of the substrate as dead weight to be optimised away.
  • 10⁻¹⁰ quoted as DNA polymerase's error rateNEGATIVE / conflation trap. It is the rate after three layers. The polymerase alone is ~10⁻⁴–10⁻⁵ in vitro. Recorded because it is the single most common error in this material.
  • 1.2 × 10⁻⁸ (per generation) set beside 10⁻¹⁰ (per replication)NEGATIVE / conflation trap. Different denominators. A generation contains many germline divisions; the numbers are not comparable and their ratio means nothing.
  • The layered stack presented as one coherent in vivo measurementNOT-RECONCILED, recorded as such. The ladder now printed (10⁻⁵ → 10⁻⁷ → 10⁻¹⁰) is Fijalkowska et al. 2012's round-decade estimate: self-consistent, but it does not reproduce from Schaaper's measured fold-discriminations. St Charles et al. 2015 measure in vivo base selectivity >100× higher than in vitro. This pass did not reconcile them and prints no reconciliation. The stacking principle survives; the rungs do not travel without their assay.
  • The power-of-ten decomposition (10⁻⁵ / 10⁻² / 10⁻³ → 10⁻¹⁰) attributed to Schaaper 1993NEGATIVE / this chapter's own misattribution, now corrected. Schaaper 1993 reports fold-discriminations (200,000–2,000,000× / 40–200× / 20–400×), not decades, and its abstract contains no "about 10⁻¹⁰" clause — a clause this chapter previously printed in quotation marks against his name. The decomposition and that clause belong to Fijalkowska, Schaaper & Jonczyk 2012, FEMS Microbiol Rev 36(6):1105–1121 — a review, citing Schaaper. A quotation mark is a claim about a source, and it was false here.
  • The ladder rung 10⁻⁴ → 10⁻⁷ under a stated gain of 10¹–10²NEGATIVE / arithmetically impossible, now corrected. 10⁻⁴ × 10⁻¹ = 10⁻⁵ and 10⁻⁴ × 10⁻² = 10⁻⁶; reaching 10⁻⁷ needs 10³, which contradicted the row's own gain, the chapter's own text, and Schaaper's measured 40–200×. The row was graded OBSERVED-REPLICATED while contradicting itself. It generated the printed ~610 figure, which survives only because the first rung moved to 10⁻⁵.
  • The 30 nm chromatin fibre drawn as an established in vivo levelNEGATIVE. Two independent near-native methods (cryo-EM of vitrified cells; ChromEMT) fail to find it. It is real in vitro. Textbook diagrams showing it inside a living nucleus are ahead of the evidence.
  • "~10.5 bp/turn" explained as a midpoint of a solution value and a surface valueNEGATIVE / this chapter's own error, now corrected. Rhodes & Klug explicitly identify their 10.6 with "the number of base pairs per turn of the DNA double helix in solution." So 10.4 and 10.6 are two assays of the same quantity under the same nominal condition, and their disagreement is method-dependent and unreconciled herenot dissolved by "two conditions, both right," which was a resolution this chapter invented and its own cited source refutes. 10.5 is printed as B-DNA's residues/turn by Potaman & Sinden Table 1; whether it is an independent determination or a rounding is NOT-SOURCED in this pass. Use it as shorthand, never as a receipt.
  • Attributing the 12/6 Å groove widths to Potaman & Sinden (NBK6545)NEGATIVE / a fabricated citation by this chapter, now corrected. NBK6545's Table 1 has no groove-width row (its rows: handedness, residues/turn, axial rise, pitch, tilt, rotation/residue, diameter, glycosidic bond, sugar pucker), and its prose treats the grooves qualitatively. The values are standard and may well be right; the receipt was invented. Recorded because of where it happened: in the very passage that lectures "name the convention or do not print the number," graded OBSERVED-REPLICATED. The chapter committed its own cardinal defect (RES IPSAE NON SIMULACRA) in the sentence warning against it. That is the most instructive failure in this file.
  • Groove widths quoted without a conventionNEGATIVE, and the convention story does not close. Both pairs are now NOT-SOURCED here. The ~5.8 Å offset accounts for the minor pair (12 → 6.2 ≈ 6) but fails on the major pair (22 → 16.2, not 12 — a 10 Å gap). "Different convention" is therefore not a sufficient explanation, and this chapter no longer asserts it as one. The ~5.8 Å figure is itself unsourced here.
  • NOT-MEASURED / NOT-SOURCED in this pass: loops, TADs, and chromosome-level organisation; the exact Franklin & Gosling 1953 pagination; both groove-width pairs (12/6 and 22/12) and the ~5.8 Å convention offset; the provenance of "10.5 bp/turn"; any reconciliation of Wang's 10.4 with Rhodes & Klug's 10.6; any reconciliation of the in vivo and in vitro fidelity stacks; a primary reporting the code-optimality figure as 10⁻⁴ directly (the primary says 0.02%).

HONEST FENCE — MODELED

Fenced MODELED. Individual rows carry their own classes (most OBSERVED-REPLICATED; several OBSERVED-CONTESTED; one INADMISSIBLE; several NOT-SOURCED). But the chapter as an artifact is a budget model: it composes measured constants through stated assumptions — a spherical nucleus, DNA as a 1 nm-radius cylinder, a mid-range 20 bp/s fork, 2 bits/bp with no sequence correlation, and a fidelity ladder assembled from different assays in different organisms. The assumptions are the fence. The origin-count prediction (~4,200) lands inside the measured band (1,000–100,000), but that band spans two orders of magnitude and is therefore a weak test — do not oversell the agreement.

Per Gould & Lewontin (1979), "The Spandrels of San Marco and the Panglossian Paradigm": nothing above establishes that any feature of DNA is an optimum. The C-value paradox is itself the loudest evidence against pan-adaptationism in this chapter — genome size is substantially drift, transposon load, and frozen accident, not design. Novozhilov et al. 2007 say the same of the code: a local optimum on a rugged landscape. Nature's authority here is precisely and only this: it already ran a very long parallel search under real physical constraints, with the failures deleted. That makes the layering principle a hypothesis generator. It does not make it a proof. Per repo rule M7, any UNI design taking the stacked-cheap-layers principle from this chapter must still beat a tuned conventional baseline (e.g. one high-quality component) on a pre-registered metric, with a discriminator that collapses the gain and a true computed-residual ablation — or it is recorded NEGATIVE.

Not claimed

  • Not claimed: that DNA is a program, code, blueprint, or specification. It is a chemically stable, high-fidelity, mechanically characterised information substrate whose readout is context-dependent and whose bulk is not specification. The metaphor's critique is cited above and is load-bearing, not decorative.
  • Not claimed: that 2 bits/bp is the genome's information content. It is an upper bound that assumes independence across positions. Real genomes are correlated; the true figure is lower and is NOT-MEASURED here.
  • Not claimed: that the non-coding 98% is junk. The C-value paradox refutes "genome size measures design content" — it does not establish what any particular sequence does. The ENCODE claim that >80% of the genome is functional is contested: Graur et al. (2013), Genome Biol Evol 5(3):578–590, doi:10.1093/gbe/evt028, argue it conflicts with the <10% estimated to be under purifying selection. This chapter takes no position and asserts neither 80% nor 10%.
  • Not claimed: that the fidelity ladder's rungs are settled. They are assay-dependent and the in vivo/in vitro discrepancy is printed, unreconciled.
  • Not claimed: that the 30 nm fibre does not exist. It exists in vitro. Its status as an in vivo structural level is contested and carried as contested.
  • Not claimed: that the genome-as-prior reading measures anything, predicts anything, or supports any UNI design. It is fenced HYPOTHESIZED and is deletable without changing one number in this chapter.
  • Not claimed: that any citation above raises any UNI rung. A nature citation is NEVER a UNI gate. The NATURA classes (OBSERVED-REPLICATED / OBSERVED-CONTESTED / MODELED / HYPOTHESIZED / INADMISSIBLE / NOT-MEASURED) and the UNI ledger's four values (proven / designed / hypothesized / not-yet-built) describe different kinds of claim and never merge. This chapter contains zero UNI claims.
  • QUAESTIO-APERTA: "full human" and "beyond human" appear nowhere here as a target, milestone, or deliverable. They are permanent open questions. Sequencing a genome has no bearing on them, and the ability to write one has no bearing on them either.

sha256 dcbec4a222de5f3c — of the original file, so what was ingested stays checkable.

Plain — written for this website, not the source document

Written for this website — not the document. This is a plain-language retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

The molecule is taken as what it physically is: a polymer with measured dimensions and mechanics, carrying a measurable amount of information at a measured error rate, packed and copied at measured rates. The molecular biology is quoted from other people's work, not produced here. None of this project's own results stand on any of it. The closing list names what is not claimed, first among them that this molecule is a program, code or blueprint.

That refusal is also the warning. Reading the molecule as source code is marked inadmissible, because no observation could refute it. The chapter's own numbers show why it misleads. A designer who believes it looks for a specification where there is only a starting bias, expects fixed behaviour where the chemistry depends on context, and treats most of the substrate as dead weight.

Its best engineering lesson is about layers. The overall accuracy is not bought by building an excellent component; it is bought by stacking several cheap, unimpressive stages whose failures are independent. That is the structural claim, and what would have to be shown to break it is named.

Several of its own earlier errors are corrected in public, including a misattributed quotation and a rung of its own ladder that was arithmetically impossible.

Plain · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is dcbec4a222de5f3c

Clear — written for this website, not the source document

Written for this website — not the document. This is a clearer retelling, written to help you meet the document. It is not the source, and it is not evidence. It has not yet been checked by a person. (or choose Precise in the reading-level control above)

This chapter is about a molecule and refuses to let the molecule become a metaphor. None of the biology here is ours: it was measured elsewhere and cited, so it is not one of the project's own results. A list at the end names what the chapter does not claim, starting with the idea that this molecule is a program or a blueprint. It opens with geometry at the resolution a designer would actually need, then treats it as a polymer with real mechanics, then gets to information content — and immediately deflates it, because that is where the expensive mistakes live.

Its structural claim is an engineering one and it is stated as such: the system-level performance comes from composing cheap layers whose failures are uncorrelated, not from any component being excellent. What would refute it is written concretely. Show a single replication component reaching the system-level accuracy on its own with the later stages disabled. Or show a large genome copied within one measured window from very few starting points, so that the throughput was bought by speed rather than by parallelism. Either would move the chapter. Individual numbers falling outside their scope move only their own rows.

The sections then walk the stack. First the error budget in layers, which the chapter calls its best engineering lesson. Then the packing hierarchy that fits a great length into a tiny volume, and the arithmetic of copying that makes a huge genome tractable at all. Last the code itself, where the redundancy is argued to be structure rather than noise. The central organising story of molecular biology is given with its genuine exceptions rather than as a slogan. A reading of all this through the book's inference framing is offered explicitly as hypothesised and as a lens only.

The recorded negatives are unusually self-implicating, and that is what makes the chapter worth reading. A famous record-holding genome size is marked inadmissible, with the defect located not in the original measurement — which was honest for its era and its method — but in continuing to cite it today. Two different error rates with different denominators are flagged as not comparable, and a widely quoted figure is shown to be the rate after several layers rather than the rate of one component. A ladder of figures is printed with an explicit note that two source measurements were not reconciled in this pass, and that no reconciliation is being offered. The stacking principle survives, but the individual figures do not travel without the assay that produced them.

Then three corrections of the chapter's own previous text. A decomposition into powers of ten was attributed to the wrong source and printed inside quotation marks, and the chapter says plainly that a quotation mark is a claim about a source and that this one was false. One step of the ladder was arithmetically impossible against its own stated gain and had been graded as if measured. And an explanation of two disagreeing values as a midpoint between different conditions was wrong, because both were assays of the same quantity under the same nominal condition, leaving a genuine method-dependent disagreement unreconciled.

A structural feature drawn in every textbook is marked negative as an in-life claim, because two independent near-native methods fail to find it — it is real in the test tube, and the diagrams are ahead of the evidence.

Clear · written 2026-08-01 by claude-opus-5 · not yet checked by a person · about the document whose sha256 is dcbec4a222de5f3c