Skip to content

25.5 — Finding the Root Cause of a Disease

For most of the twentieth century, every doctor knew what caused stomach ulcers: stress and acid. The treatment was a bland diet, antacids, and eventually surgery to cut the nerve that drives acid secretion. It was in every textbook.

In 1982 two Australians, Barry Marshall and Robin Warren, kept finding a curved bacterium in the stomach lining of ulcer patients. The stomach was supposed to be sterile — nothing could live in that acid. Their work was rejected and mocked. So in 1984 Marshall drank a culture of the organism himself, developed gastritis within days, showed the bacterium in his own stomach lining, and cured himself with antibiotics.

Ulcers are now, in most cases, an infection with Helicobacter pylori, curable in a week with two antibiotics and an acid blocker. Marshall and Warren received the Nobel Prize in 2005.

Two lessons sit inside that story and both are permanent. The first is that "we know what causes this" is very often a description of an association that everybody stopped questioning. The second is that the entire pipeline in this Part — billions of dollars, fifteen years, thousands of people — is built on top of one belief about what causes the disease. If that belief is wrong, everything downstream fails, expensively, in Phase II.

So this chapter is about how that belief gets built, and how good it has to be before anyone spends money on it.

Association is not cause, and the industry has a checklist for the difference

Almost all evidence about human disease starts as an association: people who have X more often have Y. The hard question is whether X causes Y, and there are exactly three ways an association can arise without causing anything.

Reverse causation. Y causes X, not the other way round. Early Alzheimer's disease reduces physical activity, which can make inactivity look like a cause of dementia.

Confounding. Something else causes both. The classic teaching case is that people who carry matches have higher lung cancer rates; the match is not the cause, smoking is. Confounding is the single most common reason a promising observational finding evaporates.

Chance and bias in how the data was collected. If you test enough associations, some appear by luck; if you recruit patients differently from controls, you manufacture differences.

In 1965 the British statistician Austin Bradford Hill set out the considerations that make causation likely, and they are still the working checklist. Is the association strong? Is it consistent across different populations and study designs? Is it specific to that disease? Did the cause come before the effect in time? Is there a dose–response gradient, so that more exposure means more disease? Is a biological mechanism plausible, and coherent with what else is known? Is there experimental evidence — does removing the exposure remove the disease? That last one is the strongest, and it is why a randomised trial settles arguments that decades of observation cannot.

For infections there is an older and stricter set, Koch's postulates: the organism must be found in every case, isolated and grown in pure culture, cause the disease when introduced into a healthy host, and be recoverable from that host. Marshall's self-experiment was, quite deliberately, the third postulate.

Epidemiology: finding the pattern in populations

Epidemiology is the study of who gets a disease, where, when, and what they have in common. It is where most causes are first spotted, and its vocabulary appears constantly in client documents.

Incidence is how many new cases appear in a population in a period. Prevalence is how many people have the disease at a moment. A short fatal illness has high incidence and low prevalence; a long-lasting one like diabetes has the reverse. Confusing them produces market forecasts that are wrong by an order of magnitude, which is a real and frequent error in commercial analytics work.

Three study designs cover most of the field.

The case–control study starts with people who have the disease and people who do not, and looks backwards at their exposures. It is fast and cheap, and it is the only practical design for rare diseases. Its weakness is that memory and record quality differ between the groups.

The cohort study starts with exposure and follows people forward to see who develops the disease. It is slower and far more expensive, and much stronger, because exposure was recorded before anyone knew who would get ill.

The randomised controlled trial assigns the exposure by chance. Randomisation is the only method that balances the things you did not think to measure, which is why it is the standard against which everything else is judged. Chapter 25.10 goes into how it does that.

Two numbers report the results. Relative risk is the risk in the exposed divided by the risk in the unexposed — a relative risk of 2 means twice the risk. The odds ratio is a closely related figure used in case–control studies, and it approximates relative risk when the disease is rare. Always ask what the underlying risk was. A doubling of a risk that was one in a million is not the same news as a doubling of a risk that was one in ten, and the difference between relative and absolute risk is the most common way a genuine result is reported misleadingly.

Genetics: the strongest evidence for a cause you can get in humans

Genetics has become the most valued source of causal evidence in drug discovery, for a reason that is worth understanding properly.

Your genes were assigned to you at conception, before anything in your life happened. Nothing about your diet, income, habits or environment can have altered them. So if a genetic variant is associated with a disease, reverse causation is impossible and ordinary confounding is largely ruled out. A genetic association is close to a natural experiment.

Two kinds of genetic finding get used, and they do different jobs.

Single-gene disease genes. For conditions inherited in a clear pattern, families are studied to narrow the responsible region, then the gene is identified. This is how the cystic fibrosis gene was found in 1989. Here the gene tells you exactly what the disease is, and the therapy usually aims at fixing or bypassing the broken protein.

Genome-wide association studies, or GWAS, compare many hundreds of thousands of genetic markers across large numbers of people with and without a common disease. These rarely find a single cause. They find dozens of variants, each contributing a small amount of risk, and the value is in pointing at biological pathways rather than at one gene.

The reason companies pay for this is a specific and repeatedly reproduced observation: drug targets with supporting human genetic evidence are roughly twice as likely to survive clinical development as targets without it. In an industry where 93 percent of programmes fail, doubling the odds is transformative, and it explains why large companies buy access to biobanks pairing genetic data with health records.

A related technique deserves a name because it appears constantly in evidence discussions. Mendelian randomisation uses a genetic variant that affects a modifiable factor — say, a variant that lowers lifelong cholesterol — as a stand-in for the factor itself. Because the variant was randomly inherited, comparing carriers with non-carriers approximates a lifelong randomised trial of lower cholesterol. It is how the industry gained confidence, before running the trials, that lowering a particular blood lipid would reduce heart attacks.

The omics layer, in plain terms

"Omics" is a suffix meaning "all of them at once", and the four you will hear are simple once named.

Genomics is the DNA sequence — what the cell could make. Transcriptomics is the set of RNA messages present — what the cell is currently choosing to make (Chapter 2.3). Proteomics is the proteins actually present — what is doing the work. Metabolomics is the small molecules — the chemical state the cell is in.

A modern disease study measures several of these in diseased and healthy tissue and looks for what differs. Single-cell methods now do this one cell at a time, which matters because a tumour or an inflamed joint is a mixture of very different cells whose signals cancel out when averaged.

Two honest warnings, because this is where a lot of money is spent.

Difference is not cause. Thousands of things differ in diseased tissue. Most are consequences of being ill.

And the data volume is genuinely large — a sequenced human genome is on the order of a hundred gigabytes of raw data, and studies run to thousands of samples. This is the branch of the industry where a software engineer's ordinary skills are immediately valuable: pipelines, storage, reproducibility, versioning of reference data, and the ability to re-run an analysis years later and get the same answer.

Proving it in the laboratory before betting on it

A statistical association, however strong, is not enough to start a programme. Somebody has to show that interfering with the suspected cause changes the disease, and that work happens in model systems.

Cell lines are cells grown indefinitely in flasks. Cheap, fast, and only distantly like a human tissue. Organoids are small three-dimensional clumps grown from stem cells that partly reproduce the architecture of a real organ; better, harder, newer. Genetically modified animals — usually mice — are used to remove a gene entirely, or to insert a human disease mutation, and then see what happens to the animal.

And CRISPR made a whole new approach routine: switch off every gene in the genome, one gene per group of cells, and see which switches produce the effect you want. These pooled screens are how many modern targets are found.

The permanent limitation is species. A mouse is not a small human. Immunology in particular translates badly, and a long list of treatments that cured mice did nothing for people. Chapter 25.7 covers what animal work can and cannot tell you, since that is exactly where the regulatory requirement sits.

What makes a target worth spending money on

"Target" means the specific molecule in the body — usually a protein — that the drug will act on. Target identification is finding candidates; target validation is building the case that acting on it will actually help patients.

Four questions decide it, and a client's early pipeline meetings are largely arguments about these four.

Is it causal? Does changing this molecule change the disease, in human genetics and in models, or is it merely present at the scene?

Is it druggable? Some proteins have a pocket a small molecule can occupy; some are on the cell surface where an antibody can reach them; and some — many transcription factors, many protein-to-protein contact surfaces — have historically resisted both.

Is it safe to interfere with? If the molecule does five other useful things elsewhere in the body, blocking it will produce five side effects. Human genetics helps here too: people who naturally carry a low-activity version of the protein are a preview of what your drug will do to everyone.

And can success be measured? If there is no way to tell within a reasonable time whether the drug is working, the clinical programme becomes unaffordable. That question leads directly to biomarkers.

Biomarkers, and the trap inside them

A biomarker is anything measurable that tells you something about a biological state: a blood level, a scan finding, a genetic variant. Four types get used, and the distinctions are not pedantic — regulators enforce them.

A diagnostic biomarker says whether the person has the disease. A prognostic biomarker says how their disease is likely to progress regardless of treatment. A predictive biomarker says whether this particular treatment will help them — the HER2 test from Chapter 25.3. A surrogate endpoint is a biomarker used in place of the outcome you actually care about, because it appears sooner.

The surrogate is where money and lives have been lost, and the cautionary case is famous enough to be worth carrying. Irregular heartbeats after a heart attack are associated with sudden death, so drugs that suppressed those beats were expected to save lives. The CAST trial in the late 1980s tested that assumption properly and found the treated patients died more often than those on placebo. The surrogate had moved in the right direction while the patients did worse.

So the rule the industry learned is: a surrogate is only trustworthy when there is strong evidence that changing it changes the real outcome, not merely that the two travel together. Chapter 25.10 shows how endpoints are chosen and defended, and Chapter 25.13 covers accelerated approval, which is built entirely on surrogates and carries this risk deliberately.

Where the software actually is

Everything in this chapter runs on systems somebody has to build and keep working: sequencing pipelines and their reference data, variant databases whose interpretations change over time, biobank platforms linking genetic data to health records under strict privacy rules, laboratory information systems, screening data management, and the increasingly common problem of pulling evidence out of published literature at scale.

And one requirement runs through all of it, which is the first genuinely GxP-adjacent thing you will meet: reproducibility. If a target is used in a regulatory submission years later, somebody must be able to re-run the analysis with the same code, the same reference data and the same parameters, and get the same numbers. Chapter 25.20 is where that requirement becomes law rather than good practice.

Next: Chapter 25.6, what happens once a target is chosen — how a company gets from "block this protein" to an actual molecule worth putting into a person.