Appearance
25.19 — Root Cause Analysis Done Properly
An investigation concludes: "Operator did not follow SOP-4471 step 12. Root cause: human error. Corrective action: operator retrained on SOP-4471."
That paragraph appears in thousands of quality records every year, and it is almost always wrong — not factually, but as an explanation. The operator did miss the step. The question the investigation was supposed to answer is why a trained person, doing their job, missed it, and whether the next person will too.
Ask that question and the answers are usually structural: the step is on the back of a page nobody turns, the sequence on the screen does not match the sequence at the machine, the step takes four minutes and the line waits, or three different procedures describe the same task differently. Fix any of those and the failure stops. Retrain the operator and it recurs, with a different name on it.
This chapter is the set of tools used to get from the first answer to the real one, what each is good for, and how each is misused.
What a root cause actually is
A root cause is a cause which, if removed or corrected, prevents recurrence — and which the organisation has the power to remove.
Three tests come out of that definition and they settle most arguments.
The recurrence test. If you fix this and the same event can still happen, you have not reached the cause.
The control test. "The supplier's material varied" may be true, but you cannot change another company's process by wishing. The cause you can act on is that your incoming testing did not detect the variation, or that your process has no tolerance for it.
And the honesty test. If the stated cause is a person's behaviour, ask what made that behaviour likely. Blame ends an investigation; explanation continues it.
Getting the problem statement right first
More investigations fail here than anywhere else, and it takes five minutes to avoid.
A good problem statement says what happened, to what, where, when, how much, and how it was detected — with no cause, no blame and no proposed solution in it.
Bad: "Operator error caused a temperature excursion." It names a cause before investigating, and gives no facts.
Good: "On 14 March at 03:29, the temperature in drying oven DO-3 rose to 66 degrees, 6 degrees above the upper limit of 60, for 11 minutes during batch B-2291. Detected by the alarm at 03:40. Two other batches were dried in DO-3 in the preceding week."
That statement already tells you the scope question has been asked, which is the question inspectors find missing most often.
The five tools, and when each is right
Asking why until you reach something you can fix
The 5 Whys method is simply asking why repeatedly, using the answer each time as the next question. It is fast, requires nothing, and is the right tool for a simple, single-chain problem.
Worked properly on the oven:
Why did the temperature exceed the limit? Because the heater kept running past the setpoint. Why did it keep running? Because the control valve did not close fully. Why did it not close fully? Because its seat was worn. Why was it worn without being noticed? Because the preventive maintenance schedule for DO-3 covers the heater and the fan, not the control valve. Why does the schedule omit it? Because the schedule was copied from oven DO-1, which uses a different valve type, when DO-3 was installed in 2019.
Now you have something worth fixing, and notice what it also tells you: any other equipment whose maintenance schedule was copied at installation may carry the same gap. That extension is what turns a corrective action into a preventive one.
Where the method goes wrong is well known. It assumes a single chain when reality often has several. It stops when someone reaches an answer they find comfortable. And "five" is not a rule — stop when you reach a cause you can control, whether that is at three or at seven.
Laying out every possible cause before choosing
The cause-and-effect diagram, also called a fishbone or an Ishikawa diagram after the engineer who popularised it, puts the problem at the head and groups possible causes along branches: people, method, machine, material, measurement and environment.
Its value is not the drawing. It is that it forces a group to name possibilities they would otherwise skip, and to do it before anyone has settled on a theory. Use it when the cause is genuinely unknown, when several functions are involved, and when the first suggested explanation is suspiciously convenient.
Then the discipline that makes it real: each branch is a hypothesis, and each hypothesis must be tested against evidence and marked as confirmed or ruled out, with the evidence recorded. A fishbone with twenty unexamined branches attached to an investigation report is decoration, and an inspector will read it as such.
Thinking about what could fail, before it does
Failure mode and effects analysis, FMEA, is forward-looking rather than backward-looking. For each step of a process you list how it could fail, what the effect would be, and what the cause would be. Then each failure mode is scored on three scales: how severe the effect is, how likely the cause is, and how likely you are to detect it before harm results.
Multiply the three and you get a risk priority number, used to rank where to spend effort.
Two warnings, because this tool is widely misapplied. The numbers are ordinal judgements, not measurements, so a score of 240 is not twice as bad as 120 — it is a ranking device, and treating the arithmetic as physics leads to confident nonsense. And severity must never be traded away: a failure that could kill someone is addressed regardless of how rare it is or how well you think you would detect it.
Where FMEA earns its keep is in design and in change control — deciding what a new process, a new system or a proposed change could do wrong before it does it. It is also the natural tool for deciding how much validation effort a computer system deserves, which is exactly how the risk-based approach in Chapter 25.21 works.
Following the tree of possibilities downward
Fault tree analysis starts with the undesired event at the top and works down through the combinations of conditions that could produce it, using simple logic: this failure needs A and B together, or A or B alone.
It suits complex technical systems where several things must coincide, and it makes visible the case everyone misses — the single condition that appears under several branches and therefore defeats what looked like independent safeguards.
Asking what changed
The simplest and most underused tool of all. Something worked before and does not now, so something changed.
Go through it systematically: materials and their lots, equipment and maintenance, procedures and versions, personnel and shifts, environment and season, software versions, suppliers, and volumes. In a well-run change control system (Chapter 25.17) that list is a query rather than an investigation, which is a concrete example of good records saving days of work.
Human error, treated seriously
Because "human error" is the most common recorded cause in this industry, it is worth having a real framework for it rather than a slogan.
Distinguish three different things.
A slip or lapse is doing the wrong thing while intending the right thing — clicking the adjacent field, missing a step in a familiar sequence. These are attention failures, and the fix is design: make the fields distinguishable, make the sequence enforced, make the missing step block progress.
A mistake is doing exactly what you intended, where the intention was wrong because your understanding was wrong. The fix is knowledge and clarity: better procedures, better training, clearer labelling.
A violation is knowingly departing from the procedure. Here the question is why: violations are usually routine rather than reckless, and they persist because the official method is slower, harder or impossible while production targets remain. A workaround that everybody uses is a design failure, not a discipline failure, and treating it as discipline drives it underground where it is far more dangerous.
And the hierarchy of controls says where to spend your effort, from most to least effective: eliminate the possibility; substitute something safer; engineer a control that prevents the error; use administrative measures such as procedures and checklists; and, weakest of all, warn people and rely on their care. A software system sits at the third level and can often reach the first, which is exactly why automation of a procedural control is one of the most valuable things a services company can offer a quality department.
Risk management as a formal system: ICH Q9
Everything above is investigation of things that happened. ICH Q9 is the pharmaceutical industry's framework for managing risk deliberately, and your clients will reference it constantly.
Its structure is simple. Assess risk — identify hazards, analyse their probability and severity, and evaluate them against criteria. Control risk — reduce it where needed and accept what remains, explicitly. Communicate it. And review it as new information arrives.
Two principles inside it are quoted so often that they are worth knowing verbatim in substance. The effort applied to managing a risk should be proportionate to the level of that risk. And the ultimate criterion is protection of the patient.
A 2023 revision added emphasis worth knowing about, because it addresses exactly the failure modes people complain about. It stresses that risk assessments must be based on evidence rather than opinion dressed as scoring, that subjectivity should be recognised and minimised rather than hidden behind numbers, that risk-based decisions must be documented so the reasoning survives, and that formality should itself be scaled — not every decision needs a full workshop with a scoring matrix.
That last point is genuinely useful in client conversations. A great deal of pharmaceutical risk work is theatre: elaborate matrices producing numbers that were decided in advance. The guideline itself says that is not the intent, and citing it is a legitimate way to argue for a lighter, more honest approach.
What a good investigation record looks like
Regardless of tool, the finished record must answer a defined set of questions, and if you build systems for this, these are your required fields — not because a template says so, but because an inspector will ask each one.
What happened, factually, with times and quantities. How it was detected, and how long it went undetected. What was contained, and when. What the full scope is — every batch, lot, product, patient or record that could be affected. What the cause is, with the evidence that established it and the alternatives that were ruled out. What the product impact decision is and who made it. What actions follow, split into correction, corrective and preventive, each with an owner and a date. How effectiveness will be measured, and by when. And whether anything must be reported to a regulator or a customer, decided explicitly rather than by omission.
Two more habits separate a strong quality organisation from a weak one.
Repeat-event analysis. The same failure recurring under three different record numbers is the loudest signal a quality system produces, and it is invisible unless events are coded consistently and someone looks across them. This is a data problem, and it is one of the most valuable analytics projects available in a regulated client.
And investigating near misses. An event where nothing was harmed carries the same information as one where something was, at a fraction of the cost. Organisations that investigate near misses have fewer real events, and the ones that do not are simply waiting.
Next: Chapter 25.20, the rules that decide whether any of these records can be believed at all — data integrity, ALCOA+, audit trails and 21 CFR Part 11.