Appearance
25.10 — Deciding Whether It Worked
Two groups of patients. One got the drug, one got placebo. In the drug group 18 percent had a heart attack over three years; in the placebo group 22 percent did. Did the drug work?
You cannot answer that yet, and the reasons you cannot are the whole content of this chapter. Were the groups comparable to begin with? Was the difference bigger than chance would produce? Was this the outcome the company said in advance it would judge success on, or the best-looking of forty they measured? Were patients who dropped out counted, and on which side? And is a four-point difference in that risk worth the side effects?
Every one of those questions has a formal answer built into how trials are designed, and a regulator's statistician will ask all of them.
Endpoints: choosing what counts as success, before you look
An endpoint is the measurement used to judge the trial. Choosing it is the most consequential design decision, and it must be fixed before any data is seen.
The primary endpoint is the one the trial succeeds or fails on. There is normally exactly one. Secondary endpoints support the story and can appear in the label if properly controlled statistically. Exploratory endpoints generate hypotheses and prove nothing.
Endpoints fall on a spectrum from what patients care about to what is convenient to measure.
A hard clinical outcome is an event nobody can argue about: death, a heart attack, a stroke, a hospital admission. These are what matter, and they are slow and expensive to accumulate.
A patient-reported outcome is how the patient says they feel or function, measured with a validated questionnaire. In pain, depression, or breathlessness this is not a compromise — it is the actual outcome of interest, since there is no blood test for suffering.
A surrogate endpoint is a measurement expected to predict the outcome you care about: blood pressure standing in for stroke, tumour shrinkage standing in for survival, a viral load standing in for progression to disease. They make trials smaller and faster, and they carry the risk shown by the CAST trial in Chapter 25.5, where suppressing an abnormality that predicted death produced more deaths.
Composite endpoints combine several events — say death, heart attack or urgent revascularisation — into one count, so that a trial reaches enough events to be affordable. The trap is that the components are not equal in importance, and a benefit driven entirely by the least serious component reads on a slide as a benefit on the whole composite. Asking which component drove the result is one of the most useful questions available to a non-statistician.
One more concept, recent and now unavoidable in protocols: the estimand. It forces the team to state precisely what treatment effect they are estimating, including what should happen to patients who stop treatment, switch to something else, or die of an unrelated cause. The reason it exists is that two analyses of the same trial can give different answers purely because they treated dropouts differently, and before this framework that choice was frequently made after seeing the data.
Randomisation: the only thing that balances what you did not measure
Randomisation means each patient's treatment is decided by chance rather than by anyone's judgement, and it does something no other technique can do.
Matching balances the factors you thought of. Randomisation balances everything, including factors nobody has ever named, because there is no path by which any patient characteristic can influence the assignment. That is why the randomised trial sits at the top of the evidence hierarchy, and why observational data — however large — cannot fully replace it.
Several methods are used, and their names appear in every statistical section.
Simple randomisation is a coin flip per patient. It can leave groups unequal in size by chance, particularly in small trials.
Block randomisation shuffles within small blocks so that after every block the groups are balanced. Blocks must be of varying size, otherwise a site that keeps track can predict the last assignment in each block.
Stratified randomisation runs separate schemes within categories that strongly affect outcome — disease severity, age group, country — so that those factors are balanced in each arm.
Minimisation assigns each new patient to whichever arm makes the overall balance across several factors best, with a random element retained.
And separate from the method is allocation concealment: the person enrolling the patient must not be able to discover the next assignment before deciding to enrol them. If they can, they will consciously or unconsciously steer sicker patients away from the treatment arm, and the trial is destroyed. This is why assignment is handled by the independent randomisation system described in Chapter 25.9.
Blinding, and why it is not paranoia
Blinding means people do not know who received what.
Single-blind: the patient does not know. Double-blind: neither patient nor investigator knows. Some trials also blind the people assessing outcomes and the statisticians preparing the analysis.
The reason is that knowledge changes behaviour in ways nobody intends. A patient who knows they are on the active drug reports feeling better. An investigator who knows will probe harder for improvement, and may be more willing to record an ambiguous event as unrelated. Where the endpoint requires judgement — a radiologist deciding whether a tumour grew, a neurologist scoring a movement disorder — the effect of unblinding is large and measurable.
Keeping a blind requires effort. The placebo must match the drug in appearance, taste and packaging. Where the two treatments cannot be made to look alike, trials use a double-dummy design: every patient takes both a tablet and an injection, one of which is the placebo version.
And blinds get broken accidentally by details nobody planned for. A drug that turns urine orange, a side effect only one arm produces, or a laboratory value visible to the investigator that reveals who is on treatment. This last one is a system design problem you may be asked to solve: certain results must be visible to the independent monitoring committee and hidden from the site.
Emergency unblinding must always be possible for one patient at a time, when knowing the treatment would change their medical care, and every unblinding is recorded and investigated.
Choosing the comparison
Placebo control is the cleanest scientifically and is only ethical when there is no proven effective treatment being withheld. Where an effective treatment exists, withholding it to run a cleaner trial is not acceptable, and the comparison is against that treatment — an active control.
That leads to a distinction people get wrong constantly.
A superiority trial asks: is the new treatment better than the comparator?
A non-inferiority trial asks: is the new treatment not meaningfully worse? This is the right question when the new treatment offers something other than raw efficacy — fewer side effects, an oral form instead of an infusion, no need for monitoring blood tests.
The trap in non-inferiority is the margin: the amount of inferiority you are prepared to accept, which must be justified and set in advance. Set it generously and you can declare a worse drug non-inferior. A regulator will attack that margin harder than almost anything else in the design.
The statistics you need to follow the conversation
You do not need to run these analyses. You need to understand what the numbers mean when they appear, because they drive decisions your systems support.
The p-value is the probability of seeing a difference at least this large if the treatment truly had no effect. The conventional threshold is 0.05. It is not the probability that the drug works, and it says nothing about how big the effect is — a trivial difference becomes statistically significant if the trial is large enough.
The confidence interval is more useful and increasingly preferred. A 95 percent confidence interval is the range of effects compatible with the data. "A 30 percent reduction, with an interval from 12 to 44 percent" tells you both that the effect is real and how uncertain its size is; a p-value alone tells you neither.
Power is the probability that the trial will detect a real effect of a given size. Trials are normally designed for 80 or 90 percent power. The sample size follows arithmetically from four things: the expected event rate in the control group, the size of the difference worth detecting, the power chosen, and the significance threshold. An underpowered trial is the worst kind of failure, because a negative result means nothing — you learn only that the study could not have found the answer.
Multiplicity is the problem of asking many questions. Test twenty endpoints at a threshold of 0.05 and you expect one apparent success by chance alone. Regulators therefore require a pre-specified plan for how the significance threshold is divided across endpoints, and a secondary endpoint reported without that control cannot support a claim.
Two analysis populations appear in every report. Intention to treat analyses every randomised patient in the group they were assigned to, whatever they actually took. Per protocol analyses only those who completed as intended. Intention to treat is the primary analysis for efficacy, and the reason is not pedantic: patients who stop taking a drug often stop because it was not working or was making them ill, so excluding them systematically flatters the treatment.
And missing data is the permanent, unsolved problem. Patients move, withdraw, or die. Every method of handling missing values embeds an assumption about why the data is missing, and the honest approach — now expected — is to state the assumption and show that the conclusion survives plausible alternatives, which is called a sensitivity analysis.
The statistical analysis plan, and why its timing matters
The statistical analysis plan is a document specifying every analysis in full detail: populations, models, covariates, handling of missing data, and the multiplicity strategy.
It is finalised before the database is locked and before anyone sees the treatment assignments. That timing is the whole point. Once a plan is signed, the company cannot choose the analysis that makes its drug look best, because the choice was made in ignorance of the answer. A statistical plan modified after unblinding is a red flag that any regulator will pursue.
The committee that can stop the trial
Large trials, especially those with serious outcomes, have an independent Data Safety Monitoring Board — also called a Data Monitoring Committee.
It is made up of clinicians and a statistician with no stake in the outcome. It sees the accumulating results by treatment arm, unblinded, while everyone running the trial stays blind. It meets on a schedule set in a charter written before the trial starts.
It can recommend three things. Continue. Stop for harm, when the treated group is doing worse. Or stop for overwhelming benefit, when continuing to give patients the control treatment can no longer be justified.
There is real tension in that last one. Stopping early for benefit tends to overstate the effect, because trials are more likely to be stopped at a random high point in the data, and it leaves the long-term safety picture thin. Modern charters therefore set deliberately demanding thresholds for stopping early, so that the bar reflects the permanence of the decision.
Interim analyses cost something statistically, and this surprises people. Every time you look at the data and could stop, you get another chance to be fooled by chance, so the significance threshold must be adjusted across the looks — usually by an alpha-spending function, which allocates a small part of the total error budget to each interim look.
Adaptive designs, in one paragraph each
An adaptive design allows pre-specified changes based on accumulating data, with the statistical consequences accounted for in advance.
Sample size re-estimation adjusts the number of patients if the observed variability differs from what was assumed. Dose-dropping designs start with several doses and discontinue the ones performing poorly. Seamless designs run what would have been two phases as one study, saving the gap between them. Platform trials test several treatments against one shared control group, with arms entering and leaving over time — an approach that proved its value during the COVID-19 pandemic.
The rule that makes all of it legitimate is that the adaptation must be specified before the trial starts. Changing a trial in response to results in a way that was not planned is not an adaptive design; it is data-driven decision making, and regulators treat it accordingly.
What comes out at the end
When the database is locked, the treatment codes are revealed and the pre-specified analyses run. What follows are the outputs your systems will be asked to produce, store or publish.
The tables, listings and figures — hundreds or thousands of statistical outputs, each traceable to the analysis plan and to the datasets that produced them.
The clinical study report, whose structure is defined internationally by ICH E3: the design, the conduct, every deviation, the demographics, the efficacy results, the complete safety results, and appendices reaching thousands of pages. This is the document that goes into the regulatory submission of Chapter 25.24.
Datasets in standard formats. Regulators require the analysis-ready datasets themselves, in CDISC standards — SDTM for the collected data and ADaM for the analysis datasets — along with the definitions file describing them. The agency's statisticians re-run the key analyses independently, and if the submitted data does not support the sponsor's numbers, that is discovered.
And publication. Results must be posted to the public registry, and the trial is normally published in a journal. A trial that quietly disappears because it came out badly is the practice all of this machinery was designed to end.
Next: Chapter 25.11, the physical problem nobody outside the industry thinks about — how the medicine that worked in a trial gets manufactured at a scale of millions of doses without changing.