Skip to content

9.3 — The Tests That Do Not Work, and Why They Feel Right

In 1948 a psychologist named Bertram Forer gave his students a personality test and then handed each of them what he said was their individual profile. He asked them to rate its accuracy out of five.

The average rating was 4.26.

Every student had received the identical text, assembled by Forer from a newsstand astrology book. It included lines like:

You have a great need for other people to like and admire you. You have a tendency to be critical of yourself. You have a great deal of unused capacity which you have not turned to your advantage. While you have some personality weaknesses, you are generally able to compensate for them.

This is the Barnum effect, or the Forer effect, and it has been replicated many times. Statements that are vague, generally flattering, and true of nearly everybody are experienced as strikingly personal.

Every unfalsifiable personality system runs on it, and knowing the mechanism is the only protection.

Why a wrong test feels accurate

Five mechanisms, all of them things you have already met.

The Barnum effect. The statements apply to almost everyone.

Confirmation bias (1.4). You read the description, search memory for matching instances, find several, and count that as verification. You do not search for counter-examples.

Self-fulfilment. Told you are an intuitive type, you begin to notice and report your intuitions, and to describe yourself that way to others (2.6).

Flattery. No popular type system contains an unattractive category. Every result is a strength with a mild caveat.

And the effort you put in. Answering ninety questions creates a sense of having produced something personal, which raises the credibility of whatever comes back.

None of this requires you to be gullible. Everybody does this, including psychologists who know the literature.

The Myers-Briggs Type Indicator

The most widely used personality instrument in the world, taken by millions of people a year, and it fails on three technical grounds. This matters because it is used in hiring, team-building and career advice.

It has no scientific pedigree. It was developed by Katharine Cook Briggs and her daughter Isabel Briggs Myers, neither of whom had training in psychology, based on Carl Jung's theory of psychological types — which Jung himself derived from clinical intuition rather than data, and which he explicitly cautioned against using as a classification scheme.

Its reliability is poor. Retest people after a few weeks and a substantial proportion — figures around 50 per cent within five weeks appear in the literature — come out as a different type. A measuring instrument that gives a different answer half the time is not measuring anything stable.

And the types are not real. This is the deepest problem. The MBTI sorts people into categories — you are either Thinking or Feeling — which requires the underlying dimension to be bimodal, with two clusters and a gap between them.

It is not. Every measured personality dimension is normally distributed: one hump, most people in the middle. So the cut point falls exactly where most people are, and someone scoring 51 per cent and someone scoring 49 per cent get opposite labels and are described as different kinds of person.

what types would requiretwo clusters, a genuine gapwhat is actually measuredone hump — and the cut falls where most people are
Why the type system fails at the level of arithmetic. Categories require a gap. Personality dimensions do not have one, so the boundary is drawn through the densest part of the distribution, and two people either side of it are more similar to each other than either is to someone at their own extreme.

What is genuinely useful about it. It gives people vocabulary for talking about differences, and it is non-threatening because no result is bad. Those are real benefits of a shared language, and they are not evidence that the categories exist. They also do not justify using it to make decisions about people.

Do not use it for hiring or for selecting anybody for anything. Its own publisher's guidance says so.

The Enneagram

Nine types, each with wings, levels of health and stress and growth directions.

Origins: a twentieth-century synthesis drawing on the teachings of George Gurdjieff and later Oscar Ichazo and Claudio Naranjo, presented in some accounts as having ancient roots that historians do not support.

Evidence: thin. The type structure has not been established through the kind of factor analysis that produced the Big Five, and reliability and validity studies are limited and mixed.

Why it feels deeper than MBTI. Because it addresses motivation and fear rather than behaviour, and because the descriptions include unflattering material — which defeats the flattery mechanism and therefore feels braver and more accurate.

Its honest value: the questions it asks are good ones. What are you most afraid of, what do you most want, what do you do under stress. Those are worth answering. The nine boxes are not required in order to answer them.

The rest

Colour systems used in corporate training — red, blue, green, yellow. Simplified type systems with the same problems and less pedigree.

Astrology. No mechanism, and studies testing whether astrologers can match birth charts to personality profiles find performance at chance. Its persistence is a demonstration of the Barnum effect at scale.

Blood-type personality, widely believed in parts of East Asia. No supporting evidence.

Handwriting analysis. Repeatedly tested, does not predict personality. Still occasionally used in hiring, which is worse than useless because it introduces noise.

And the Rorschach inkblot test. More complicated than the others: a subject of serious ongoing dispute within psychology, with some scoring systems having modest validity for a small number of specific things and the broad personality interpretations having little. Not a fraud, and not what popular culture thinks it is.

What to use instead

The Big Five, from a research instrument. Dimensions rather than types, decent reliability, and it actually predicts things.

But hold even that lightly. The correlations are modest (9.2), self-report is imperfect (4.5), and a profile describes tendencies rather than a person.

And for practical self-knowledge, a record beats any test. Two weeks of logging what you actually did and how you actually felt (4.5) tells you more about yourself than any questionnaire, because it measures behaviour rather than your beliefs about your behaviour.

Why this matters beyond being right

Three real costs.

Decisions get made. People are hired, assigned, promoted and advised on the basis of type results. A test with poor reliability used for selection is injecting noise into people's careers.

It becomes an identity, and identities constrain (8.7). I'm an introvert, so I don't do networking. I'm a Type Four, so I'm meant to be melancholic. A label acquired from an invalid test becomes a reason not to attempt things, and the reason is not real.

And it crowds out the useful question. "What type am I?" is a question with no true answer. "What do I actually do, in which situations, at what cost?" is answerable, and the answer is useful.

What to do with this

Take the Barnum text at the top of this page seriously as a test of yourself. If it felt somewhat accurate — and it does to most people — you have just experienced the mechanism from the inside, which is worth more than reading about it.

Use the Big Five if you use anything. Free, dimensional, and it predicts.

And convert any type description you find compelling into a testable statement. I am an introvert becomes I am more tired after three hours of company than most people I know. The second one can be checked, and checking it is self-knowledge while accepting the first is a label.

Next: 9.4 takes the most-used of those labels — introversion — and separates it properly from shyness, sensitivity and social anxiety, which are four different things routinely treated as one.