Appearance
2.1 — DNA: the Molecule, and How It Was Found
Every cell in your body carries the same two metres of DNA, wound up inside a nucleus six micrometres across. It holds about 3.1 billion letters. If you read those letters aloud at one per second, without sleeping, it would take just under a century.
That molecule specifies every protein you will ever make. This page is about what it physically is, and about the twenty years of argument it took to work that out — because the argument is where the reasoning lives, and the reasoning is what makes the structure make sense.
Nobody believed DNA was the genetic material
For the first half of the twentieth century, almost everyone assumed genes were made of protein. The logic was reasonable. Proteins are built from twenty different amino acids and can be arranged in an unimaginable number of ways. DNA has only four different units. A four-letter alphabet looked far too simple to carry the instructions for building an organism, and DNA was widely dismissed as structural packing material.
Three experiments changed that.
Griffith, 1928. Frederick Griffith was working with two strains of pneumococcus bacteria. The smooth strain has a sugar coat, causes fatal pneumonia in mice, and looks glossy on a plate. The rough strain has no coat, is harmless, and looks dull. He injected mice with four preparations. Living rough bacteria: mice lived. Living smooth: mice died. Heat-killed smooth: mice lived. Heat-killed smooth mixed with living rough: the mice died, and Griffith recovered living smooth bacteria from their bodies.
Something from the dead smooth bacteria had permanently converted the harmless rough strain into the deadly smooth one, and the change was inherited by their descendants. Griffith called it the transforming principle. He had no idea what it was.
Avery, MacLeod and McCarty, 1944. They spent a decade purifying that transforming principle and testing it by elimination. They destroyed the protein with protein-digesting enzymes — transformation still worked. They destroyed RNA — still worked. They destroyed DNA with DNase — transformation stopped.
That should have settled it, and it did not, because the assumption that genes were protein was so entrenched that most biologists concluded the DNA preparation must be contaminated with a trace of some special protein.
Hershey and Chase, 1952. The experiment that finally convinced everyone used a virus that infects bacteria — a bacteriophage, the kind pictured in Chapter 1.1. A phage is essentially a protein shell containing DNA, and it works by injecting something into a bacterium that makes the bacterium produce more phages. The question was: which part goes in?
They grew one batch of phages with radioactive sulfur and another with radioactive phosphorus. The choice of elements is the whole experiment. Protein contains sulfur, in the amino acids cysteine and methionine, and contains no phosphorus. DNA contains phosphorus, in its phosphate backbone, and contains no sulfur. So each label tags one component and only that component.
They let each batch infect bacteria, then spun the mixture in a kitchen blender to shear the empty phage shells off the bacterial surfaces, and centrifuged so the heavy bacteria went to the bottom and the light shells stayed in the liquid. The radioactive phosphorus was inside the bacteria. The radioactive sulfur was in the liquid with the discarded shells.
The DNA went in. The protein stayed outside. The argument was over.
The clues that were on the table by 1952
Once DNA was accepted as the genetic material, three separate pieces of evidence existed about its structure, and the person who assembled them would have the answer.
The chemistry. DNA is a chain of nucleotides, each made of deoxyribose sugar, a phosphate, and one of four bases: adenine (A), guanine (G), cytosine (C), thymine (T). A and G are purines, built on a double ring. C and T are pyrimidines, built on a single ring. So a purine is physically about twice the width of a pyrimidine — a fact that turns out to matter enormously.
Chargaff's rules, 1950. Erwin Chargaff measured base composition across many species and found something strange. The proportions differ between species — human DNA is about 30 percent A, E. coli about 26 percent — but within any one species, the amount of A always equals the amount of T, and the amount of G always equals the amount of C. Nobody knew why. It looked like a numerical coincidence.
Photograph 51, 1952. Rosalind Franklin, working at King's College London, produced X-ray diffraction images of DNA fibres of a quality nobody had achieved. When X-rays pass through a regular structure they scatter in a pattern determined by that structure, and the pattern in Franklin's best image was a bold X shape. An X-shaped diffraction pattern is the signature of a helix, and the spacings in it gave the helix's dimensions: a repeat every 3.4 nanometres, with a layer every 0.34 nanometres, meaning ten units per turn, and an overall diameter of about 2 nanometres. Franklin had also determined that the phosphate backbone must be on the outside.
The image was shown to James Watson by Maurice Wilkins, Franklin's colleague, without her knowledge, and her unpublished report was seen by Watson and Crick through a Medical Research Council committee. This was the central injustice of the story. Franklin died of ovarian cancer in 1958 at 37, four years before the 1962 Nobel Prize went to Watson, Crick and Wilkins, and Nobel Prizes are not awarded posthumously — but the credit problem existed long before her death and was not primarily about the prize rules.
The structure

Watson and Crick published a one-page paper in Nature in April 1953. Here is the structure, feature by feature, with what each feature explains.
Two chains, wound into a double helix. Diameter 2 nanometres, one complete turn every 3.4 nanometres, ten base pairs per turn — exactly Franklin's numbers.
The sugar–phosphate backbones are on the outside, the bases on the inside. The backbone is charged and water-friendly, so it faces the water; the bases are flat and largely water-avoiding, so they stack in the middle like coins in a roll. This is the same hydrophobic effect that builds the membrane in Chapter 1.4, operating inside one molecule.
A pairs only with T. G pairs only with C. This is the key that unlocked everything. A purine (wide) always pairs with a pyrimidine (narrow), so every rung is the same width, and the helix has a constant diameter — which the diffraction pattern demanded. Pair two purines and the rung would bulge; pair two pyrimidines and it would pinch.
And it explains Chargaff instantly. If every A is opposite a T, the totals must be equal. The coincidence was not a coincidence; it was a fingerprint of the pairing rule.
The pairs are held by hydrogen bonds — the weak links from Chapter 1.2. A–T is held by two, G–C by three. So a stretch rich in G and C is harder to pull apart than a stretch rich in A and T, and this is not a curiosity: it is why a laboratory heats DNA to separate the strands at a temperature that depends on its composition, and it underlies every PCR test you have ever had.
The two strands run in opposite directions — antiparallel. Each strand has a chemical direction, conventionally called 5′ (five prime) to 3′ (three prime), named for which carbon atom of the sugar carries the next link. One strand runs 5′→3′ upward, its partner runs 5′→3′ downward. This looks like a detail. It is not, and Chapter 2.2 shows why it forces the copying machinery into an awkward and revealing compromise.
The strands are complementary, not identical. If one reads GATTACA, the other must read CTAATGT. So each strand contains all the information needed to rebuild the other.
That last point is what Watson and Crick pointed at in the paper's famously restrained closing line: "It has not escaped our notice that the specific pairing we have postulated immediately suggests a possible copying mechanism for the genetic material." Separate the two strands, and each one is a template for a new partner. The structure did not merely fit the data — it explained heredity, which no protein model had ever come close to doing.
Reading the molecule's own dimensions
Some arithmetic worth doing yourself, because it makes the scale real.
Your genome is about 3.1 billion base pairs per copy, and each base pair adds 0.34 nanometres of length.
3.1\times10^{9} \times 0.34\times10^{-9}\ \text{m} \approx 1.05\ \text{m}
That is one copy. You have two — one set from each parent — so about two metres of DNA per cell, packed into a nucleus six micrometres wide. The packing ratio is roughly 200,000 to 1. Chapter 2.7 explains how it is done without knotting.
And across your body: two metres times 30 trillion cells is about 6 × 10¹³ metres of DNA, which is roughly 400 times the distance from the Earth to the Sun.
What "a gene" actually is
A gene is a stretch of DNA that specifies a product — usually a protein, sometimes a functional RNA. Humans have roughly 19,000 to 20,000 protein-coding genes, a number that shocked people when the Human Genome Project reported it, since estimates before sequencing had run to 100,000. A roundworm has about 20,000 too.
Gene count is not complexity, and Chapter 2.4 explains why: what matters is regulation and the ability to make several different proteins from one gene.
Only about 1 to 2 percent of your DNA codes for protein. The rest was once dismissed as junk. Much of it genuinely does nothing — old virus sequences and broken copies of genes make up a large share — but a substantial part is regulatory, controlling when and where genes are switched on. Deciding how much is functional remains an active argument, and anyone who states a confident single figure is overselling.
An allele is a version of a gene. You have two copies of most genes, one from each parent, and they may be the same version or different ones. That distinction is the whole of Chapter 2.6.
The genome is the complete set of DNA in one cell. Your nuclear genome is 3.1 billion base pairs; your mitochondrial genome, from Chapter 1.5, is 16,569.
Where this touches you already
A PCR test works because of the hydrogen bonds in this page. Heating the sample to about 95 °C breaks the hydrogen bonds and separates the strands. Cooling to around 55 to 60 °C lets short designed pieces of DNA — primers — find and stick to their exact complementary sequence, and only that sequence, because base pairing is specific. An enzyme then extends from each primer. Repeat the cycle 30 to 40 times and each cycle doubles the amount, so 2³⁰ is roughly a billion-fold amplification from a starting quantity too small to detect. Chapter 16.4 covers what that means for diagnosis.
A DNA vaccine or an mRNA vaccine is this molecule used as a message. Chapter 13.5.
And every inherited disease in this volume is a change in this sequence — sometimes one letter out of three billion, as in sickle cell (Chapter 2.8).
What the next page fixes
Watson and Crick's closing sentence proposed a copying mechanism and left it there. Doing it in practice turns out to be far harder than the sentence suggests, because the two strands run in opposite directions and the copying enzyme can only work in one of them. Chapter 2.2 follows a replication fork and shows how the cell solves that problem, how it achieves an error rate of about one in a billion, and what happens to you when the proofreading fails.