Skip to content

2.2 — Replication: Copying Three Billion Letters Without Ruining Them

Every time one of your cells divides, it copies 3.1 billion base pairs, twice over, in about eight hours. The final error rate is roughly one mistake per billion letters copied — about three errors per complete genome copy.

To see how extraordinary that is: a professional typist copying text makes an error every few hundred characters. The cell is a million times more accurate, and it is doing it at about fifty letters per second per site, on a molecule that is a metre long and must not be tangled.

It manages this with three layers of accuracy, stacked. This page follows the copying machine and then takes each layer apart.

Semi-conservative: proving how it copies before knowing how it works

Watson and Crick's structure suggested that each strand serves as a template for a new partner, giving two daughter molecules that each contain one old strand and one new one. This is called semi-conservative replication — half of each product is conserved from the parent.

Two rivals were possible. Conservative: the parent double helix stays intact and a completely new one is built. Dispersive: the copies are patchworks of old and new segments.

Matthew Meselson and Franklin Stahl settled it in 1958 with an experiment often called the most beautiful in biology, and the beauty is that it distinguishes all three possibilities in one run.

They grew bacteria for many generations in a medium where the only nitrogen was the heavy isotope ¹⁵N, so every base in their DNA was slightly heavier than normal. Then they moved the bacteria to ordinary ¹⁴N medium, so all new DNA would be light, and sampled after each round of division. Each sample was spun in a caesium chloride gradient, which sorts molecules into sharp bands by density.

After one generation there was a single band, exactly halfway between heavy and light. That immediately killed the conservative model, which predicts one heavy band and one light band with nothing in between.

After two generations there were two bands: one intermediate, one fully light. That killed the dispersive model, which predicts a single band getting steadily lighter and never separating.

Semi-conservative predicts exactly what was seen: after one round, every molecule is one heavy strand plus one light strand — all intermediate. After two rounds, the intermediate molecules split, and half the products are intermediate while half are fully light. Two bands, in equal amounts.

The machinery at a fork

Diagram of a replication fork showing helicase unwinding the double helix, primase laying down RNA primers, DNA polymerase extending the leading strand continuously and the lagging strand in fragments, and ligase joining the fragments
A replication fork. Helicase splits the helix at the point of the fork. One new strand (the leading strand) is built continuously toward the fork; the other (lagging) is built backwards in short fragments that are later joined. Both new strands grow in the same chemical direction — that is the constraint that forces the awkward arrangement. Image: Wikimedia Commons.

Origins. Copying does not start at one end. It starts at specific sequences called origins of replication — one in a bacterium, tens of thousands scattered along human chromosomes, because a single fork moving at 50 base pairs per second would take about two years to copy a human chromosome. With many origins working simultaneously the whole genome is done in hours.

Helicase unwinds the double helix, breaking the hydrogen bonds and separating the strands. It moves along ahead of everything else, at about 1,000 base pairs per second in bacteria.

Single-strand binding proteins coat the separated strands to stop them snapping back together or folding on themselves.

Topoisomerase solves a problem that is easy to overlook. Unwinding a twisted rope in the middle makes the rest of the rope twist tighter ahead of you. In a molecule a metre long, that supercoiling would seize up within seconds. Topoisomerase cuts one or both strands, lets the tension unwind, and reseals. This enzyme is a major drug target: the quinolone antibiotics such as ciprofloxacin block the bacterial version, and several chemotherapy drugs including etoposide and doxorubicin block the human version, killing dividing cells by leaving their DNA cut.

Primase lays down a short RNA primer. This is needed because of a hard limitation of the main copying enzyme: DNA polymerase can only extend an existing chain — it cannot start one. It needs a free 3′ end to add to. Primase, which makes RNA, has no such restriction, so the cell starts every new stretch with a scrap of RNA and replaces it later.

DNA polymerase does the copying. It reads the template and adds the complementary nucleotide: opposite an A it places a T, opposite a G it places a C. And it can only build in the 5′→3′ direction, adding to the 3′ end.

The lagging strand problem

That last restriction, combined with the antiparallel structure from Chapter 2.1, creates the central awkwardness of replication.

The fork opens in one direction. On one template strand, building 5′→3′ happens to point toward the fork, so the polymerase simply follows the opening fork continuously. This is the leading strand, and it is copied in one smooth run.

On the other template, building 5′→3′ points away from the fork. The polymerase cannot run backwards. So it does something clumsy: it waits until helicase has opened a stretch, then runs backwards along that stretch away from the fork, stops, and waits for the next stretch. This is the lagging strand, and it is built as a series of short pieces called Okazaki fragments — about 150 to 200 bases each in humans, named after Reiji and Tsuneko Okazaki, who found them in 1968.

Each fragment starts with its own RNA primer. So afterwards the cell must:

  1. Remove every RNA primer (an enzyme with RNase activity does this).
  2. Fill the gaps with DNA (a repair polymerase).
  3. Seal the nicks between fragments (DNA ligase, which forms the final bond).

The lagging strand is not a design flaw so much as an unavoidable consequence of two facts — the strands run antiparallel, and the chemistry of adding a nucleotide only works at one end. Every organism on Earth has the same problem and the same solution, which tells you it was already present in LUCA.

Three layers of accuracy

Layer 1: base pairing itself. The shapes and hydrogen bonding patterns mean the correct base fits far better than a wrong one. On its own this gives an error rate around 1 in 10,000 — nowhere near good enough.

Layer 2: proofreading. DNA polymerase has a second active site that runs backwards. After adding a base it checks the fit; if the new pair is distorted, the polymerase reverses, cuts the wrong nucleotide out, and tries again. This improves accuracy about a hundredfold, to roughly 1 in a million.

Layer 3: mismatch repair. A separate system sweeps behind the fork looking for the small distortion a mismatched pair leaves in the helix. When it finds one, it cuts out a stretch of the new strand around the error and resynthesises it. Another thousandfold improvement, reaching about 1 in a billion.

The obvious question is how the repair system knows which strand is new — cutting out the original would convert a repairable mismatch into a permanent mutation. In bacteria the answer is chemical marking: the parent strand carries methyl groups added some minutes after synthesis, so for a brief window the new strand is unmethylated and identifiable. In humans the mechanism is different and relies partly on the nicks still present in the newly made strand, and it is not fully understood.

When mismatch repair fails, you get a specific cancer syndrome. Lynch syndrome — hereditary non-polyposis colorectal cancer — is caused by an inherited fault in one of the mismatch repair genes. Mutations accumulate across the genome at maybe a hundred times the normal rate, and affected people have a lifetime colorectal cancer risk of roughly 50 to 80 percent, often striking before 50, along with raised risk of endometrial and other cancers. It accounts for perhaps 3 percent of all colorectal cancers. Chapter 19.5.

And that same defect is now a treatment signal. Tumours with broken mismatch repair accumulate enormous numbers of mutations, which means they display enormous numbers of abnormal proteins on their surface — so the immune system can recognise them if it is released to do so. Immunotherapy works unusually well in these tumours for exactly that reason, and mismatch repair status is now tested routinely to decide treatment. Chapter 19.7.

The end of the chromosome

There is one place where replication cannot finish, and it is a direct consequence of the primer requirement.

On the lagging strand, the very last Okazaki fragment near the chromosome's end starts with an RNA primer. When that primer is removed there is no upstream 3′ end for a polymerase to extend from, so the gap cannot be filled. Every round of replication therefore loses a short stretch from each chromosome end — roughly 50 to 200 base pairs in human cells.

This is the end-replication problem, and it is why telomeres exist: expendable repeated sequence at the ends, so that the loss eats buffer rather than genes. Chapter 1.7 covered the consequences — the Hayflick limit, senescence, and telomerase in cancer.

What breaks it, and what that gives medicine

Every step above is a target, and several of the most important drug classes in medicine sit here.

Quinolone antibiotics (ciprofloxacin, levofloxacin) block bacterial topoisomerase. Bacterial and human topoisomerases differ enough for this to be selective. Chapter 22.6.

Chemotherapy drugs of several classes attack replication in human cells and therefore hit whatever is dividing fastest. Etoposide and doxorubicin trap human topoisomerase in the middle of its cut. Methotrexate and 5-fluorouracil starve the cell of the nucleotide building blocks. Cisplatin chemically links the two strands so they cannot separate at all. Chapter 19.7.

Antiviral drugs exploit the fact that some viruses carry their own polymerase. Aciclovir, used for herpes and shingles, is a nucleoside that only becomes active after a viral enzyme modifies it — so it is activated almost exclusively inside infected cells — and once incorporated into a growing viral DNA chain it stops further extension, because it lacks the 3′ attachment point the next nucleotide needs. That is why aciclovir is remarkably safe compared with most antivirals: an uninfected cell never activates it. Chapter 22.6.

Zidovudine and the other HIV drugs of its class work by the same chain-termination principle against HIV's reverse transcriptase. Chapter 17.9.

What the next page fixes

You now have a molecule that can be copied faithfully. Copying it is not using it. Chapter 2.3 covers how a sequence of bases becomes a sequence of amino acids — transcription into RNA, the genetic code read triplet by triplet, and translation on the ribosome — and why a single deleted letter can destroy a protein completely while a substituted letter often changes nothing at all.