Skip to content

1.3 — Proteins and Enzymes

Crack an egg into a hot pan. The white goes from clear and runny to opaque and solid in about thirty seconds, and no amount of cooling will ever bring it back. Nothing was added, nothing was burnt, and the chemical formula of what is in the pan is the same as what came out of the shell. The atoms did not change. The shape did, and that was enough to turn a liquid into a solid permanently.

That is the whole subject of this page. A protein is a chain of amino acids that has folded into one specific three-dimensional shape, the shape is what does the job, and losing the shape loses the job forever.

Four levels of structure

Four panels showing primary structure as a chain of beads, secondary as a coiled helix and a pleated sheet, tertiary as a folded three-dimensional blob, and quaternary as several such blobs assembled together
The four levels of protein structure. Each one is built out of the level before it: a sequence coils into local patterns, the patterns fold into a whole shape, and several whole shapes may lock together into a working unit. Image: Wikimedia Commons.

Primary structure is the sequence — which amino acid comes first, second, third, all the way to the end. Nothing else. A protein of 300 amino acids has a primary structure that is just a list of 300 items chosen from twenty. This is the level that DNA specifies directly, and everything above it follows automatically from it.

The number of possible sequences is worth pausing on. For a chain of only 100 amino acids there are 20¹⁰⁰ possible sequences, which is about 10¹³⁰ — a number vastly larger than the count of atoms in the observable universe. Life uses a vanishingly small, carefully selected corner of that space.

How much does one position matter? Take haemoglobin, the protein that carries oxygen in your red blood cells. Its beta chain is 146 amino acids long. At position 6 there is normally glutamic acid, which carries a negative charge and is therefore comfortable on the water-facing outside of the protein. Change that single position to valine, which is hydrophobic, and you have created a small greasy patch on the surface of a molecule that should have none. Greasy patches stick to other greasy patches. When oxygen levels drop, these altered haemoglobin molecules polymerise into long stiff rods that deform the whole red blood cell into a crescent — a sickle. The stiff cells jam in small vessels, tissue downstream dies, and the person has an episode of agonising pain called a sickle cell crisis.

One amino acid out of 146, in a chain built from a DNA sequence where one letter changed. That is the entire cause of sickle cell disease, and it is the cleanest demonstration in medicine that sequence is destiny. Chapter 2.8 follows it all the way back to the DNA.

Secondary structure is local pattern. As the chain sits in water, the backbone starts hydrogen-bonding to itself — specifically, the N—H of one amino acid to the C=O of another a few places away. Two arrangements do this especially well and therefore appear everywhere.

The alpha helix is a right-handed coil, about 3.6 amino acids per turn, with every backbone N—H hydrogen-bonded to the C=O four positions back down the chain. Linus Pauling worked it out in 1948 while ill in bed in Oxford, folding paper models, before anyone had solved a protein structure. Keratin — hair, nails, the outer layer of skin — is mostly alpha helix.

The beta pleated sheet is the other: sections of chain lie side by side, either running the same direction or opposite, hydrogen-bonded across the gap, forming a slightly corrugated sheet. Silk is almost entirely beta sheet, which is why it is strong in tension and does not stretch.

Note what is doing the work here: hydrogen bonds between backbone atoms, not side chains. That is why these two patterns are universal — the backbone is the same in every protein.

Tertiary structure is the whole shape, and this is where the side chains take over. Once the local helices and sheets exist, the chain folds back on itself so that:

  • Hydrophobic side chains get buried in the middle, away from water. This is the dominant force, and it is not really an attraction between the greasy parts — it is water pushing them together so it can go back to hydrogen-bonding with itself. This is the same effect that makes oil droplets merge in a glass of water, operating inside a single molecule.
  • Charged side chains pair up — a positive lysine finds a negative aspartate and they hold each other with an ionic bond.
  • Polar side chains hydrogen-bond, both to each other and to the surrounding water on the outside.
  • Two cysteines can form a disulfide bridge, an actual covalent bond between their —SH groups. This is much stronger than everything else on this list and acts as a rivet, locking part of the shape permanently.

The result is a specific object with pockets, grooves and protruding loops. That object's surface is the protein's function. A pocket exactly the shape of glucose is a glucose-binding site. A groove exactly the shape of a DNA double helix is a DNA-binding site.

Quaternary structure exists only in proteins built from more than one chain. Haemoglobin is the standard example: four separate folded chains — two alpha, two beta — clicked together into one unit, each holding an iron atom that binds one oxygen molecule. The four cooperate, which is a crucial property covered in Chapter 8.4: when one subunit picks up oxygen it shifts shape slightly and makes the other three bind more eagerly. That cooperation is what lets your blood load oxygen almost fully in the lungs and unload a large fraction of it in a working muscle.

Who tells the chain how to fold?

Nobody. That is the remarkable part.

Christian Anfinsen tested this directly around 1961 with an enzyme called ribonuclease. He unfolded it completely with urea and a reducing agent, destroying every disulfide bridge and every fold, until it was a random chain with no activity at all. Then he removed the chemicals. The chain refolded by itself into exactly the original shape, with all four disulfide bridges reconnecting to the correct partners, and full activity returned. No template, no machinery, no instructions beyond the sequence itself.

The conclusion, which won him the Nobel Prize in 1972, is that the amino acid sequence contains all the information needed to specify the folded structure. The folded shape is simply the lowest-energy arrangement that sequence can reach, and the chain falls into it the way a dropped chain falls into a heap — except that this heap is the same heap every time.

There is a genuine puzzle sitting inside that, named after Cyrus Levinthal, who pointed it out in 1969. If a protein of 100 amino acids had to search all its possible shapes to find the best one, and could test one every picosecond, the search would take longer than the age of the universe. Yet real proteins fold in microseconds to seconds. So folding cannot be a search. It is more like water running downhill: the energy landscape is funnel-shaped, so almost every path leads downward toward the same destination, and the chain never has to try more than a tiny fraction of the possibilities.

In a real cell, folding gets help — not with the instructions, but with the crowding. The inside of a cell is dense, and a half-folded chain with its greasy parts still exposed is at risk of sticking to the nearest other half-folded chain instead of to itself. Chaperones are proteins whose job is to give a folding chain a protected chamber to fold in and to give misfolded chains another attempt. Some of them are called heat-shock proteins, because cells make far more of them when heat starts destabilising everything.

Predicting the shape from the sequence was an open problem for fifty years, and it is now largely solved. DeepMind's AlphaFold2, evaluated at the CASP14 assessment in 2020 and published in 2021, predicts protein structures at an accuracy comparable to laboratory measurement for a large fraction of proteins. Structures for essentially every known protein have since been released publicly. David Baker, Demis Hassabis and John Jumper shared the 2024 Nobel Prize in Chemistry for this work. It is one of the few places where the machine learning of Volume I, Part 12 has directly changed what is possible in medicine, because knowing a target's shape is the first step in designing a drug that fits it.

When folding goes wrong

A folded protein rendered in three dimensions, showing ribbon-like helices and arrow-like sheets packed against each other into a compact globular shape
A folded protein rendered from its measured atomic coordinates. Ribbons are alpha helices, flat arrows are beta strands, and the loops between them are the connections. The compact packing is not decoration — it is what puts specific atoms in specific places to make a working site. Image: Wikimedia Commons.

Misfolding is not a rare accident. It is behind a whole class of disease, and the pattern repeats.

Cystic fibrosis, in about 70 percent of cases, is caused by a mutation called F508del that deletes a single amino acid from a chloride channel protein. The protein is otherwise functional, but it now folds slightly wrong, and the cell's quality control machinery spots the error and destroys it before it ever reaches the membrane. The channel is not broken; it never arrives. Chapter 2.8 covers what follows.

Alzheimer's disease involves a fragment called amyloid-beta that misfolds into beta sheets and stacks into plaques between neurons, alongside a protein called tau that misfolds into tangles inside them. Chapter 20.4 covers the disease and, importantly, the honest state of the argument about which of these is cause and which is consequence.

Parkinson's disease involves alpha-synuclein clumping inside neurons of a small brain region called the substantia nigra.

Prion diseases are the extreme version, described in Chapter 1.1: a misfolded protein that converts the correctly folded version on contact, so the misfolding itself is infectious.

The common thread is that beta sheets stack. A misfolded protein exposing beta-sheet edges will hydrogen-bond to the next one, and the resulting fibre is chemically extremely stable — which is exactly why the body cannot clear it.

Denaturation: the egg in the pan

Denaturation is loss of the folded shape while the sequence stays intact. Every force holding tertiary structure together — hydrogen bonds, ionic pairs, hydrophobic packing — is weak, so it does not take much.

Heat makes molecules vibrate harder until the weak bonds shake apart. This is the egg. Human enzymes work best near 37 °C, begin to lose activity above about 40 °C, and denature seriously above roughly 42 °C. This is why a fever above 42 °C is a medical emergency in itself, quite apart from whatever caused it — it is not the infection killing you at that point, it is your own proteins coming undone. Chapter 13.4 covers why a fever is useful at lower temperatures and where the useful range ends.

Acid or alkali adds or removes charges on the side chains, so ionic pairs break and the shape collapses. Your stomach uses this deliberately: stomach acid at around pH 1.5 to 2 denatures the proteins in your food, unfolding them so digestive enzymes can reach the peptide bonds inside. The same effect is why lemon juice or vinegar "cooks" fish in ceviche without heat, and why milk curdles when it sours.

Alcohol disrupts hydrogen bonding and pulls hydrophobic cores apart. This is how alcohol hand rub kills bacteria and viruses: it denatures their proteins wholesale. Interestingly, 70 percent alcohol works better than 100 percent, because the water is needed to carry the alcohol into the microbe rather than instantly fixing its surface.

Heavy metals like lead and mercury bind sulfur atoms in cysteine side chains and wreck the disulfide bridges. That is the basic mechanism of heavy metal poisoning.

Denaturation is usually irreversible in practice, because in a crowded environment an unfolded chain finds a neighbour to stick to before it finds its own correct shape. Anfinsen's dilute test tube was a special case.

Enzymes: proteins that make reactions fast enough to matter

Most of the reactions your body depends on would happen anyway, given a few centuries. That is not useful. An enzyme is a protein that speeds up a specific reaction, by factors that are hard to take seriously until you see them.

Carbonic anhydrase converts carbon dioxide and water into carbonic acid — the reaction that lets your blood carry CO₂ back from tissues to your lungs. A single molecule of it handles up to about a million reactions per second. Without it the reaction is so slow that you would not be able to clear CO₂ fast enough to survive.

What a catalyst actually does

Energy diagram with reactants on the left and products on the right, showing a tall hill between them for the uncatalysed reaction and a much lower hill for the catalysed one
Why a catalyst works. The reaction goes downhill overall in both cases — the products sit lower than the reactants either way. What the catalyst changes is the height of the hill in between. A lower hill means far more molecules have enough energy to get over it at body temperature. Image: Wikimedia Commons.

Two molecules that could react still have to collide with enough energy to break their existing bonds before the new ones form. That energy requirement is the activation energy, drawn as the hill in the figure. At any temperature, only a small fraction of molecules have that much energy at a given moment, and that fraction is what sets the reaction rate.

An enzyme lowers the hill. It holds the reacting molecules in exactly the right orientation, strains the bonds that need to break, and provides charged groups that stabilise the awkward halfway state. It does not change where the reaction ends up — the energy difference between start and finish is untouched, so an enzyme cannot make an uphill reaction go downhill. It only changes how fast equilibrium is reached. And the enzyme comes out unchanged, ready for the next one.

The size of the effect: lowering the activation energy by about 6 kilojoules per mole roughly multiplies the rate by ten at body temperature. Enzymes routinely lower it by ten times that.

The active site, and how specific it is

Diagram showing a substrate approaching an enzyme, the enzyme's active site changing shape to close around it, the reaction occurring, and the products leaving
Induced fit. The substrate approaches, the active site closes around it and changes shape as it does so, the reaction happens, and the products leave — releasing the enzyme unchanged. The older "lock and key" picture had the site already the right shape; induced fit is what actually happens, and the closing motion is part of the catalysis. Image: Wikimedia Commons.

The active site is a pocket on the enzyme's surface, usually made of side chains from parts of the chain that are far apart in sequence and only brought together by folding. The molecule it acts on is the substrate.

Emil Fischer proposed in 1894 that the fit was like a lock and key — rigid and exactly complementary. Daniel Koshland corrected this in 1958 with induced fit: the site is not pre-shaped, it moulds itself around the substrate as it binds, and the act of moulding is part of what strains the substrate's bonds and makes them break.

Specificity is extreme, and this is what makes enzymes useful rather than merely fast. An enzyme that acts on glucose will typically ignore galactose, which differs only in the orientation of one hydroxyl group. Many enzymes distinguish left-handed from right-handed versions of the same molecule.

Naming is mostly regular. Enzymes end in -ase, attached to what they act on or what they do: lactase splits lactose, protease splits protein, lipase splits lipids, DNA polymerase builds DNA, ATP synthase makes ATP. The few that break the rule are old names kept out of habit — pepsin, trypsin, chymotrypsin, all digestive proteases.

How fast, and why it stops getting faster

Put a fixed amount of enzyme in a tube and add more and more substrate. At first the rate climbs in proportion: more substrate means more collisions with active sites. Then the climb flattens, and eventually adding more substrate does nothing at all. Every active site is occupied, and the enzyme is working as fast as it can turn over. That ceiling is called Vmax.

Leonor Michaelis and Maud Menten described this in 1913 with an equation that still carries their names:

v = \frac{V_{\max}\,[S]}{K_m + [S]}

Read it aloud: the rate v equals V-max times the substrate concentration, divided by K-m plus the substrate concentration. Two things to notice, and both are just arithmetic.

When [S] is much smaller than K_m, the bottom of the fraction is roughly K_m alone, so v \approx (V_{\max}/K_m)[S] — the rate is proportional to substrate, which is the initial straight climb.

When [S] is much larger than K_m, the bottom is roughly [S], which cancels the [S] on top, leaving v \approx V_{\max} — the flat ceiling.

And when [S] equals K_m exactly, the fraction becomes V_{\max}/2. So K_m is the substrate concentration at which the enzyme runs at half its maximum speed, and it is a practical measure of how tightly an enzyme grips its substrate. A low K_m means the enzyme is saturated even when substrate is scarce, so it works well at low concentrations.

That number has a clinical use. Your liver has two enzymes that process alcohol, with very different K_m values. The main one, alcohol dehydrogenase, saturates at low blood alcohol levels — which is why alcohol is cleared at an almost fixed rate, roughly one standard drink per hour, regardless of how much you drank. The clearance does not speed up when you drink more, because the enzyme is already flat out.

What changes an enzyme's speed

Temperature raises the rate until it destroys the enzyme. Rates roughly double for every 10 °C up to the optimum, then fall off a cliff as the protein denatures. Human enzymes peak near 37 °C, which is not a coincidence — it is why your body temperature is what it is.

pH matters because the charges on the active site's side chains depend on it. Each enzyme has an optimum, and the optima are set by where the enzyme has to work. Pepsin, which works in the stomach, peaks around pH 1.5 to 2 and is destroyed at neutral pH. Trypsin, which works in the small intestine after the pancreas has neutralised the acid, peaks around pH 8. Salivary amylase peaks near pH 7 and stops the moment the food you swallowed reaches your stomach acid.

Cofactors and coenzymes. Many enzymes are inactive on their own and need a helper. A cofactor is usually a metal ion — zinc, magnesium, iron, copper — sitting in the active site doing chemistry the amino acids cannot. A coenzyme is a small organic molecule that shuttles something in or out.

Here is the part that connects to your dinner: most coenzymes are made from vitamins. NAD⁺ comes from niacin, vitamin B3. FAD comes from riboflavin, B2. Coenzyme A comes from pantothenic acid, B5. Thiamine pyrophosphate comes from B1.

This is what a vitamin is. A vitamin is not a fuel and not a building material — it is a small molecule your body cannot make, which is needed in tiny amounts because it is used catalytically over and over. And it is why vitamin deficiencies produce such specific and dramatic diseases: knock out one coenzyme and one exact step of metabolism stops while everything upstream of it keeps running. Beriberi from thiamine deficiency and pellagra from niacin deficiency are both this, and both are in Chapter 24.2.

Inhibitors, which is to say: most of your medicine cabinet

An inhibitor is anything that slows an enzyme down. There are two basic ways.

Competitive inhibition. The inhibitor looks enough like the substrate to sit in the active site, but nothing happens to it. It simply occupies the site and blocks the real substrate. Because it is a competition, adding more substrate wins it back — the effect can be overwhelmed.

Non-competitive inhibition. The inhibitor binds somewhere else on the enzyme and changes its shape, so the active site no longer works. Adding more substrate does not help, because the problem is not occupancy.

Both are how drugs work, and the list is long.

  • Statins competitively inhibit HMG-CoA reductase, the rate-controlling enzyme in your liver's cholesterol production line. Blocking it lowers LDL cholesterol. Chapter 22.7.
  • Aspirin inhibits cyclooxygenase, the enzyme that makes prostaglandins, which are the molecules that produce pain, fever and inflammation. Aspirin's version is irreversible — it chemically attaches an acetyl group to the enzyme and permanently disables it, which is why its effect on platelets lasts for the platelet's whole 8-to-10-day life and why you stop it before surgery. Chapter 22.5.
  • ACE inhibitors, the blood pressure drugs whose names end in -pril, block angiotensin-converting enzyme and interrupt the hormone loop that raises blood pressure. Chapter 22.7.
  • Penicillin inhibits the bacterial enzyme that cross-links cell wall material. The bacterium keeps growing, cannot reinforce its wall, and bursts. You have no such enzyme and no such wall, which is exactly why penicillin can be given at doses that destroy bacteria and leave you unharmed. Chapter 22.6.

And one case where the drug is the poison's twin. Methanol — wood alcohol, occasionally present in illegally distilled liquor — is not very toxic itself. The damage is done by what your liver turns it into: alcohol dehydrogenase converts it to formaldehyde and then to formic acid, which destroys the optic nerve and causes permanent blindness, and then kills. The treatment is to stop that conversion. Fomepizole blocks the enzyme directly. If fomepizole is not available, the emergency treatment is to give the patient ethanol — ordinary drinking alcohol — because ethanol is the enzyme's preferred substrate and competitively occupies it, so the methanol passes out in the urine unconverted. A doctor deliberately intoxicating a poisoned patient is textbook competitive inhibition, and Chapter 23.6 covers the protocol.

Enzymes as a window into your own body

Enzymes are supposed to be inside cells. When cells are damaged, they leak, and measuring them in blood tells a doctor which cells are dying.

ALT and AST are liver enzymes, and raised levels mean liver cells are being injured — by a virus, alcohol, a drug, or fat. Creatine kinase rises when muscle is damaged, including heart muscle, and very high levels appear in rhabdomyolysis, where crushed or overworked muscle floods the blood with its contents and can block the kidneys. Amylase and lipase rise sharply in acute pancreatitis, where the pancreas begins digesting itself with its own enzymes.

These three lines on a blood report are read in exactly this way every day in every hospital, and Chapter 16.5 puts numbers to them.

What the next page fixes

Enzymes only work if the right molecules are present at the right concentration in the right compartment. That requires a boundary that lets some things through and stops others, and the ability to pump things against their natural direction. Chapter 1.4 builds the cell membrane out of the phospholipids of Chapter 1.2, and shows how a cell controls what crosses it — including the pump that consumes a fifth of all the energy you spend at rest.