Skip to content

2.3 — From Gene to Protein: Transcription, the Code, and Translation

A gene is a sequence of DNA bases. A protein is a sequence of amino acids. Getting from one to the other requires a dictionary, and that dictionary is the same in bacteria, in oak trees, in fungi and in you — which is why a bacterium can read a human gene and manufacture human insulin.

This page covers the whole route: DNA copied into RNA, RNA edited, RNA read three letters at a time on a ribosome, and a protein coming off the end. Francis Crick called the overall direction the central dogma: information flows DNA → RNA → protein, and not back from protein to nucleic acid.

Step 1: transcription

Diagram of transcription showing RNA polymerase moving along DNA, unwinding a short stretch, and building an RNA strand complementary to the template strand
Transcription in progress. RNA polymerase opens a short bubble in the DNA, reads one strand, and builds a complementary RNA copy. The DNA closes again behind it, so only about seventeen base pairs are ever open at once. Image: Wikimedia Commons.

RNA polymerase binds a region just upstream of the gene called the promoter, opens a short bubble in the double helix, and reads one of the two strands. It builds an RNA copy of the other strand, adding nucleotides in the 5′→3′ direction just as DNA polymerase does.

Three differences from replication are worth stating explicitly, because they explain a lot downstream.

RNA uses uracil instead of thymine. Wherever the template has an A, the RNA gets a U rather than a T. Uracil is chemically simpler and cheaper to make. DNA uses thymine instead because of a repair issue covered in Chapter 2.5 — it makes a particular kind of damage detectable.

RNA uses ribose instead of deoxyribose, which has one more oxygen. That extra oxygen makes RNA more chemically reactive and much less stable — an RNA molecule falls apart in hours where DNA lasts for millennia. This is a feature, not a defect. A message you want to switch off should not persist, and it is also why mRNA vaccines need deep-freeze storage and lipid packaging while DNA vaccines do not.

Only one strand is read, and which one depends on the gene. Different genes along the same chromosome are read from different strands.

Where it stops: in humans, a specific sequence signals the end and the RNA is cut loose.

The edit: introns, exons and splicing

In bacteria, the RNA that comes off the gene is ready to use. In humans it is not, and the reason is one of the genuine surprises of twentieth-century biology.

Human genes are interrupted. The coding sequence is broken into pieces called exons (expressed), separated by stretches called introns (intervening) that are transcribed and then cut out. A typical human gene is about 27,000 bases long but yields a message of only about 1,300 bases — more than 95 percent of the transcript is discarded. The dystrophin gene, the largest in the human genome, is 2.4 million bases long and takes about sixteen hours to transcribe, for a message of 14,000 bases.

Splicing is done by a machine called the spliceosome, which recognises the sequences at intron boundaries, cuts there, and joins the exons.

Three modifications complete the message:

  • A 5′ cap — a modified guanine stuck on the front, which protects the message from being chewed up and tells the ribosome where to start.
  • A poly-A tail — a run of 50 to 250 adenines added to the back. This is a timer: the tail is nibbled away over time, and when it is gone the message is destroyed. A long tail means a long-lived message.
  • Splicing as above.

And splicing is not fixed. The same transcript can be spliced different ways in different tissues, including or skipping particular exons, so one gene can produce many different proteins. This is called alternative splicing, and it happens to the majority of human genes. It is the main answer to the puzzle in Chapter 2.1 — how 20,000 genes build something as complex as a human when a roundworm has the same number. We do not have more genes. We use each one in more ways.

Splicing errors cause disease directly. Some forms of beta-thalassaemia are caused by a mutation that creates a false splice site, so the message is cut in the wrong place and the resulting protein is useless. Chapter 2.8.

Step 2: the genetic code

Circular chart of the genetic code showing all 64 three-letter RNA codons and the amino acid each specifies, with start and stop codons marked
The genetic code. Read from the centre outward: the first base, then the second, then the third gives the amino acid. Notice how often the third base makes no difference — that redundancy is a built-in error tolerance. Image: Wikimedia Commons.

The code is read in triplets, called codons. Each codon of three RNA bases specifies one amino acid.

Why three? Pure arithmetic, and it is worth doing. There are four bases and twenty amino acids to specify. One base per amino acid gives 4 possibilities — not enough. Two bases give 4² = 16 — still not enough. Three bases give 4³ = 64, which is more than enough. Three is the smallest number that works, and evolution took it.

Sixty-four codons for twenty amino acids means the code is redundant: most amino acids have several codons. Leucine and serine have six each; methionine and tryptophan have only one.

Three properties matter clinically.

AUG is the start codon, and it also codes for methionine, so every protein begins life with a methionine at the front (usually removed afterwards).

UAA, UAG and UGA are stop codons. They specify no amino acid; they mean "release the finished protein". A mutation that creates one of these in the middle of a gene truncates the protein, and Chapter 2.5 covers how destructive that is.

The redundancy is not random. Where an amino acid has several codons, they usually differ only in the third base. So a mutation in the third position of a codon very often changes nothing at all. This is called a silent mutation, and it means the code has error tolerance built into its structure. Where the third base does matter, the alternatives usually specify a chemically similar amino acid, so even a real change tends to be a mild one.

The code is nearly universal. The same 64 codons mean the same 20 amino acids in E. coli, in yeast, in a whale and in you — which is the strongest single piece of evidence for common descent (Chapter 1.9) and the reason genetic engineering works at all. There are a handful of exceptions, the most important being that human mitochondria use a slightly different code, with UGA meaning tryptophan rather than stop.

Step 3: translation

Diagram of translation showing messenger RNA threading through a ribosome with transfer RNA molecules bringing amino acids and a growing protein chain emerging
Translation. The messenger RNA threads through the ribosome; each transfer RNA arrives carrying one amino acid and pairs its three-base anticodon with the matching codon; the growing chain is passed from one transfer RNA to the next and emerges from the ribosome as a protein. Image: Wikimedia Commons.

The problem translation solves is that a codon has no chemical affinity for its amino acid. Nothing about the shape of GCA attracts alanine. An adapter is needed, and Crick predicted one before it was found.

Transfer RNA (tRNA) is that adapter. Each tRNA is a small folded RNA with two business ends: an anticodon of three bases at one end, and an attachment site for one specific amino acid at the other. A tRNA with the anticodon CGU carries alanine, and only alanine.

The accuracy lives in a separate enzyme. A family of enzymes called aminoacyl-tRNA synthetases — one for each amino acid — attaches the right amino acid to the right tRNA. They proofread, and their error rate is around 1 in 10,000. The ribosome itself does not check whether the amino acid on a tRNA is correct; it only checks the codon–anticodon pairing. So if a synthetase makes a mistake, the wrong amino acid goes in and nothing catches it.

The ribosome has three slots for tRNA, and translation is a three-beat cycle repeated once per amino acid:

  1. A site (arrival) — a tRNA whose anticodon matches the next codon docks here.
  2. P site (peptide) — the growing protein chain is transferred from the tRNA in the P site onto the new amino acid in the A site, forming a peptide bond. The catalysis is done by RNA, not protein — the ribosome is a ribozyme, as Chapter 1.9 covered.
  3. E site (exit) — the ribosome shifts by exactly three bases; the now-empty tRNA moves here and leaves; the chain moves to the P site; the A site is open for the next one.

Repeat until a stop codon arrives in the A site. No tRNA matches it, a release factor binds instead, and the finished protein is let go.

Speed: about 20 amino acids per second in human cells, roughly 20 per second in bacteria too. A 300-amino-acid protein takes about fifteen seconds.

And a message is read many times at once. Multiple ribosomes queue along one mRNA, each at a different point, forming a polysome. One message can be producing dozens of copies of its protein simultaneously.

What happens after: folding and delivery

The chain that emerges is not yet a protein in the functional sense. It folds — usually starting before it has even finished emerging, sometimes with chaperone help (Chapter 1.3). It may then be modified: sugars added, phosphates attached, sections cut out.

Insulin is the standard example of cutting. It is made as a single chain called preproinsulin. A signal sequence at the front is removed as it enters the endoplasmic reticulum, giving proinsulin. Then a middle section — the C-peptide — is cut out entirely, leaving two chains held together by disulfide bridges. That is insulin.

And the discarded C-peptide is clinically useful. It is released in equal amounts to insulin, so measuring it tells a doctor how much insulin the person's own pancreas is making. Injected insulin contains no C-peptide. So a diabetic patient with high insulin and no C-peptide is injecting it; high insulin with high C-peptide means their own pancreas is producing it — which is how a rare insulin-producing tumour is distinguished from surreptitious insulin use. Chapter 18.7.

Where this is attacked in medicine

Nearly every step is a drug target, and most of them are antibiotics, for the reason given in Chapter 1.5: bacterial and human ribosomes are different enough to attack separately.

  • Rifampicin blocks bacterial RNA polymerase. A first-line tuberculosis drug (Chapter 17.8).
  • Tetracyclines block the tRNA from docking in the A site of the bacterial 30S subunit.
  • Aminoglycosides (gentamicin, streptomycin) bind the 30S and cause misreading, so the bacterium builds nonsense proteins.
  • Macrolides (azithromycin, erythromycin) and chloramphenicol block the 50S subunit and stop the chain being extended.
  • Diphtheria toxin and ricin both attack human translation directly, which is why both are so lethal in tiny amounts.

And one drug class works in the other direction. Nonsense-readthrough drugs and antisense oligonucleotides are designed to fix a broken message rather than block a working one. Nusinersen, for spinal muscular atrophy, is a short piece of engineered nucleic acid that binds a specific splice site on a backup gene's transcript and forces the spliceosome to include an exon it normally skips — turning a nearly useless backup gene into a working one. It transformed the outlook for an illness that used to kill most affected infants before two. Chapter 21.8.

What the next page fixes

Every cell in your body contains the same genes, yet a liver cell makes albumin and a pancreatic beta cell makes insulin and neither makes the other. Nothing on this page explains that. Chapter 2.4 covers gene regulation: how a cell decides which genes to read, how that decision is remembered through cell division, and how the environment leaves marks on your DNA that can outlast you.