Appearance
1.3 — How a Claim Is Tested, and the Crisis That Followed
In 2010 a study appeared in a good journal showing that people who stood in an expansive "power pose" for two minutes before a stressful task felt more powerful, took more risks, and showed measurable changes in two hormones. It became one of the most watched talks in the world. Schools taught it. Job applicants did it in toilets before interviews.
Then other laboratories tried to reproduce it. The hormone changes did not appear. The risk-taking did not appear. A large replication with several times the original sample size found nothing but a small subjective feeling of confidence, which nobody doubted in the first place. One of the original authors publicly stated that she no longer believed the effect was real.
Nothing dishonest happened. Everyone involved was a competent scientist following the normal rules of the field. The normal rules were the problem, and this page is about what those rules are, how they failed, and what changed — because if you understand this page you can evaluate a psychological claim yourself for the rest of your life, and if you skip it you will spend that life believing whichever study got the best headline.
The shape of an experiment
Strip away the detail and almost every experiment in this volume has the same skeleton.
A question with two possible worlds. Does writing about your worries before an exam improve the score? World one: it does nothing. World two: it does something.
A manipulation. You do the thing to one group and not to another. The group that gets nothing, or gets a neutral activity of the same length, is the control group, and it exists because a great many things improve on their own. People get better at tasks with practice, illnesses resolve, moods lift. Without a control you cannot tell your treatment from the passage of time.
Random assignment. Which group each person lands in is decided by chance, not by preference or by the researcher. This is the single most powerful device in the whole method and it deserves a moment. If people choose their own group, the groups differ before you start — the ones who volunteer for the writing exercise are probably more anxious, or more conscientious, or both. Randomise, and every difference between the groups, including the ones you never thought to measure, is scattered evenly by chance. Any difference at the end is then attributable to the one thing you controlled. Without randomisation you have a correlation. With it, you have a cause.
Blinding where possible. The participant should not know which group they are in; ideally the person measuring the outcome should not either. Both know how to unconsciously produce the expected result.
An outcome you decided on in advance. This turns out to be the load-bearing word, and the next section explains why.
What a p-value actually says, and what everyone thinks it says
You run the experiment. The writing group averages 71 and the control group averages 67. Is that real, or is it the kind of gap you would get from two random groups doing nothing in particular?
Statistics answers a very specific and rather narrow question here. The p-value is the probability of getting a difference at least this large if there were genuinely no effect at all. By convention, if that probability is below 0.05 — less than one in twenty — the result is called statistically significant and gets published.
Three things about that convention, all of which matter.
It does not tell you the probability that your claim is true. People read "p = 0.03" as "97 per cent likely to be real". It is not. It is a statement about how surprising your data would be in a world with no effect, which is a different sentence entirely. Whether the claim is likely true also depends on how plausible it was to begin with, and a study showing that people can feel events before they happen with p = 0.01 is far more likely to be a fluke than a discovery.
One in twenty is not rare. Run twenty experiments on effects that do not exist, and on average one comes out significant. Now consider that thousands of researchers are running experiments, that the ones which find nothing are much harder to publish, and you have a machine that fills journals with flukes. This is the file drawer problem: the negative results are sitting in file drawers, so the published record is a biased sample of what was actually found.
And the threshold is arbitrary. 0.05 was a rule of thumb suggested by the statistician Ronald Fisher in the 1920s. There is no natural law at one in twenty. A result at p = 0.049 and a result at p = 0.051 are essentially the same result, and the fact that one gets published and the other does not is an accident of history that shaped a century of science.
The number that matters more: effect size
Statistical significance answers "is there probably something there?" It says nothing about how much. With a big enough sample, an effect so small it could not possibly matter to anybody will be significant.
The measure to look for is effect size, and the common one is Cohen's d, which expresses the gap between two groups in units of how much people vary anyway.
d = \frac{\bar{x}_1 - \bar{x}_2}{s}
Read that aloud: dee equals the difference between the two group averages, divided by the typical spread within a group. The bar over the x means average, and s is the standard deviation, which is roughly the average distance of a person from their own group's mean. Dividing by it is what makes the number comparable across studies measuring completely different things.
The rough interpretation, which comes from Jacob Cohen and is a convention rather than a law:
| Cohen's d | Called | What it looks like in life |
|---|---|---|
| 0.2 | Small | Groups overlap almost completely |
| 0.5 | Medium | A careful observer would notice |
| 0.8 | Large | Obvious without measuring |
| 1.5+ | Very large | Rare in psychology outside the obvious |
Some anchors, so the numbers mean something. The height difference between adult men and women is about d = 1.7 — enormous, and yet you still meet plenty of women taller than plenty of men. The effect of spaced practice over cramming on later recall is often d = 0.7 to 1.0, which is why Part 3 treats it as the most valuable single technique in learning. The effect of most one-off "mindset" interventions is d = 0.1 or below.
A useful habit: convert d back into overlap. At d = 0.5, if you pick one person at random from each group, the person from the higher-scoring group is on top about 64 per cent of the time. Not 100. Sixty-four. Almost every psychological group difference you will ever read about looks like that, which is the statistical reason why using a group finding to predict an individual is a mistake.
The four ways a result becomes a mirage
None of these require dishonesty. That is what makes them dangerous.
Small samples. With twenty people per group, the estimate of the effect bounces around wildly. Small studies do not just find effects less reliably — when they do find one, it is systematically too big, because only an exaggerated estimate clears the significance bar at that sample size. This is why an exciting first result so often shrinks in later, larger studies.
Flexibility during analysis. The researcher has dozens of small choices: which outliers to remove, whether to control for age, which of four outcome measures to report, whether to stop collecting data now or after fifty more participants. Make each choice in the direction that helps, and you can produce a significant result from pure noise with alarming ease. In 2011 three researchers demonstrated this by using ordinary, defensible analysis choices to "prove" that listening to a Beatles song made people eighteen months younger. This is p-hacking, and most of it is done by people who believe they are being reasonable.
Deciding the hypothesis after seeing the data. You measured twelve things, one came out significant, and you write the paper as though that was the prediction all along. It sounds harmless. It converts a one-in-twenty fluke into a confident finding, and it removes the entire logic that made the p-value mean anything.
Publication bias. Journals want positive, surprising, clean results. So a real effect of zero, studied by fifty labs, produces two or three published papers all showing an effect, and forty-seven unpublished null results nobody sees.
What happened when the field checked
Starting around 2011, psychologists began systematically re-running famous experiments.
The largest single effort, the Open Science Collaboration, published in 2015, repeated 100 studies from three top journals with larger samples. Where 97 of the originals had reported significant effects, 36 of the replications did, and where an effect did survive, its size was on average about half the original. Other large replication projects in psychology and in adjacent fields have landed in a broadly similar range.
Some famous casualties, named because you will meet them being quoted as fact:
Ego depletion — the idea that willpower is a limited resource that runs down like a fuel tank across the day. A pre-registered multi-lab replication with over two thousand participants found essentially nothing. Part 8.6 covers what replaced it.
Power posing, as above: the hormonal and behavioural claims did not survive.
Social priming — that reading words about old age makes you walk more slowly, and similar effects. Repeatedly failed to replicate.
The facial feedback effect in its strong form — that holding a pen in your teeth to force a smile makes cartoons funnier. A seventeen-lab replication found nothing; later work with different methods finds something very small, and the argument continues.
The Stanford prison "experiment" was never a controlled experiment at all. Recordings released later show the guards were explicitly coached towards harshness and one famous breakdown was, by the participant's own account, performed. 11.2 tells the whole story, because the correct lesson from it is nearly the opposite of the one everybody learned.
What held up perfectly well. The Stroop effect. The spacing effect and the testing effect (all of Part 3). Classical and operant conditioning. The basic memory findings. Cognitive behavioural therapy's effectiveness for depression and anxiety. The Big Five personality structure. Most work with large effects and simple measurements survived; the casualties were concentrated among small-sample studies of surprising, subtle social effects.
What changed, and why the field is now better than its reputation
Pre-registration. Researchers now commonly publish their hypothesis, sample size and analysis plan before collecting data, in a timestamped public record. This removes the flexibility problem at a stroke. When effects are tested this way, they routinely come out smaller, and a lot of them come out at zero.
Registered reports. A journal accepts the paper based on the plan, before results exist, and publishes whatever comes out. Positive results in registered reports run at roughly half the rate of conventional papers — which tells you exactly how much the old system was selecting on outcome.
Bigger samples, and many labs at once. Studies with thousands of participants across dozens of countries are now common where a few dozen students used to do.
Open data. The raw numbers are posted so anybody can check the analysis.
This is a field that publicly audited itself, found a serious problem, and rebuilt its methods. Very few fields have done that. When you see a psychology finding, the most useful single question is now: was this published before or after about 2015, and was it pre-registered?
What to do with this
Six checks, in order, for any claim about the mind — in a newspaper, a bestseller, a video, or this book.
One: is there a control group? No comparison, no finding. This alone eliminates most claims made for wellness programmes, courses and apps.
Two: how many people? Under about fifty per group for a subtle social effect, treat it as a hypothesis, not a result.
Three: what is the effect size? If the article gives only a p-value or only a percentage improvement, be suspicious. Ask what d was, and if the answer is 0.15, the thing is real in the way a single raindrop is wet.
Four: has it been replicated by people who were not hoping it would work? One study is a suggestion. Three independent replications is a finding.
Five: does the headline say "causes" when the study only had a correlation? "People who meditate are less anxious" and "meditating reduces anxiety" are completely different claims, and only a randomised trial supports the second. Calm people may simply be more likely to take up meditation.
Six: how surprising is it? A very surprising claim needs much stronger evidence than an unsurprising one, because the base rate of surprising claims being true is low. This is not conservatism. It is arithmetic.
And apply it to yourself, because you run the same failure modes. You believe things about your own mind — that you work better under pressure, that you cannot function without eight hours, that a certain person always ruins your mood — that were formed from a handful of memorable instances with no control condition and no record. The remedy is the same one the field adopted: write down the prediction in advance, then collect data. Two weeks of a one-line daily note will overturn at least one thing you were sure of about yourself. That is not a metaphor for science. It is the same procedure.
Next: 1.4 turns this from a method for reading studies into a method for checking your own thoughts before you act on them.