Appearance
25.17 — The Quality System: Deviations, CAPA, Change Control and Complaints
At 03:40, an operator notices the temperature in a tablet-drying oven touched 6 degrees above its allowed limit for eleven minutes before the alarm brought her over. The batch is worth about four hundred thousand dollars.
What happens next is the same in every regulated company in the world, and if you can follow it end to end you can follow ninety percent of the conversations inside a quality department.
She stops what she can stop and secures the batch so nobody moves it. She records what she saw, when, and what she did, before the end of her shift. Within a day the event is entered into the quality system as a deviation. Someone assesses the risk: could this affect the product, and could it affect anything else? An investigation establishes why the oven overshot. If the cause could recur, a corrective and preventive action is opened. Every batch that could have been affected is identified. And nobody releases that batch to patients until the whole chain is closed and a quality reviewer signs it.
Those four record types — deviation, investigation, CAPA and change control, plus complaints coming in from outside — are the working machinery of quality. This chapter is what each one is, how they connect, and the traps that make them fail.
The vocabulary, settled once
These words are used loosely in conversation and precisely in systems, and the precision matters because each type has its own workflow and its own clock.
A deviation is a departure from an approved instruction, procedure or specification. The oven excursion is one. So is performing a step in the wrong order, or missing a scheduled check.
A planned deviation is a deliberate, pre-approved, one-time departure — a documented decision to do something differently for a specific batch, with justification and quality approval in advance. A large number of planned deviations against the same procedure is itself a finding, because it means the procedure no longer describes reality and should be changed instead.
A nonconformance is a product, material or result that does not meet its requirement. Tablets outside the weight range, a delivered material failing incoming inspection, a result outside specification. The distinction people struggle with is simple in practice: a deviation is about the process, a nonconformance is about the thing produced.
An out-of-specification result is a laboratory result outside its acceptance criteria, and it carries the special investigation rules from Chapter 25.12. An out-of-trend result is inside specification but moving in a direction that predicts trouble — a signal worth acting on before it becomes a failure, and the reason trending exists at all.
A complaint is any communication from outside the company alleging a deficiency in a product's identity, quality, durability, reliability, safety or performance.
And a CAPA — corrective and preventive action — is the formal project opened to eliminate the cause of a problem so it does not recur.
What actually happens when an event occurs
The sequence below is common to essentially every regulated company. Learn it once and you can read any quality system's data model.
Detect and record. The event is written down promptly by the person who saw it, in their own words, with times and facts rather than conclusions. "Temperature reached 66 degrees for eleven minutes" is a fact; "operator error" is a conclusion nobody has earned yet.
Contain. Stop the harm spreading. Quarantine affected material, stop the line, secure records, suspend distribution of anything at risk.
Classify by risk. Usually critical, major or minor. Critical means patient safety or product quality is likely affected, or a regulatory requirement was breached; major means a significant departure with no direct patient impact; minor means a documentation-level issue with no product effect. The classification decides the depth of the investigation and how far up the organisation it must be notified, so it is the first genuinely consequential decision and it is frequently argued over.
Assess the impact — and this is where inexperienced investigations fail. The question is never only "is this batch affected". It is: which other batches ran through the same equipment, which other products used the same material lot, has this occurred before, and is product already out in the market? An investigation that answers only about the batch in front of you is the single most common inspection finding worldwide.
Investigate the cause. Covered properly in Chapter 25.19, since the tools deserve their own chapter.
Decide on the product. Release, release with restrictions, rework where permitted, or reject. Only the quality unit can make this decision, and it must be justified in writing against the evidence.
Decide on actions. A correction fixes the immediate problem — repair the oven. A corrective action removes the cause so it cannot recur — replace a failing control valve and adjust the alarm setpoint so it triggers before the limit is breached. A preventive action addresses a cause that has not yet produced a failure — checking the other three ovens for the same worn valve.
Close, with evidence. Every action completed, documented, and — for CAPAs — verified as effective.
The CAPA lifecycle, and the two ways it goes wrong
A CAPA runs through a defined lifecycle: initiation, investigation, action plan with owners and due dates, implementation, effectiveness check, and closure.
The effectiveness check is the part everyone underestimates and inspectors go straight to. It asks, some defined time after implementation, whether the problem actually stopped happening — measured, not asserted. "Procedure updated and staff retrained" is not evidence of effectiveness; "zero recurrences in the following six months across 47 batches, against 9 events in the prior six months" is.
Now the two failure modes, because both are extremely common and both are recognisable from the outside.
The first is retraining as a reflex. When the cause is recorded as human error and the action is retraining, almost nothing has been fixed. People do not fail randomly; they fail where a process makes failure easy. An inspector seeing many CAPAs whose action is retraining concludes the company is not finding real causes, and they are usually right. This is also the single largest opportunity a software engineer has in a quality context: a system that makes the wrong action impossible removes the failure permanently, which no amount of training does.
The second is CAPA overload. Open a CAPA for everything and you get hundreds of open records, overdue actions, and a backlog nobody can work through. Overdue CAPAs are one of the most reliable predictors of a bad inspection, because they show a company that identifies problems and does not resolve them. Mature companies therefore use risk to decide what becomes a CAPA and handle minor issues as corrections with trending — which is a defensible position when the reasoning is documented, and an indefensible one when it is not.
Change control: the rule that no regulated thing changes silently
Change control is the process for evaluating, approving, implementing and verifying any change to a validated or regulated item — a procedure, a piece of equipment, a material, a supplier, a facility, a computer system, or a manufacturing process.
The steps are always the same. A change is proposed with a reason. It is assessed for impact: on product quality, on validation status, on other processes, on documentation and training, and — critically — on the regulatory filing, since some changes require notification to or approval from a regulator before they may take effect (Chapter 25.24). It is approved by everyone whose area is touched, including quality. It is implemented with everything that goes with it: updated documents, retraining, revalidation. And it is verified and closed.
Two variants exist because reality demands them. A temporary change is time-limited and must be reversed or made permanent by a stated date. An emergency change may be implemented before the full paperwork is complete — but only with quality approval at the time, and with documentation completed immediately afterwards. Emergency changes are scrutinised heavily, because "emergency" is an easy label to abuse.
For a software engineer this is the concept that most needs translating from your world into theirs. Your deployment pipeline is a change control system already — it just needs to produce evidence a regulator would accept: what changed, why, who approved it, what was assessed, what was tested, who verified it worked, and how it could be reversed. Chapter 25.21 makes that mapping explicit.
Complaints: the outside world reporting back
A complaint is the only quality signal that comes from actual use, which makes it the most valuable and the hardest to handle, because it arrives incomplete, in ordinary language, and often through a call centre.
The handling process is defined and each step matters.
Intake, with enough detail to be useful: the product, the batch or serial number, what happened, when, to whom, whether anyone was harmed, and whether the complainant still has the item.
Immediate triage for the two questions with clocks attached. Was there an adverse event, meaning did anyone suffer harm? If so, it enters the pharmacovigilance path of Chapter 25.22 in parallel, with its own reporting deadlines. And for devices, is it reportable to the regulator under the reporting rules of Chapter 25.16? These two determinations must be made quickly, because the deadlines run from when the company first became aware, not from when someone got around to assessing it.
Investigation. Where possible the returned sample is examined, retained samples of the batch are tested, and the batch record is reviewed. Where the item is not returned — the usual case — the investigation must reach a conclusion anyway, and "could not be confirmed" is an acceptable conclusion only when the search was genuine and documented.
Response to the complainant, and trending. The trend is where the value is: one report of a cracked container is noise, eleven reports from one production month is a problem, and only a system that codes complaints consistently can tell the difference.
And there is a specific American obligation for drugs worth knowing by name. A Field Alert Report must be submitted for an approved drug when significant chemical, physical or other change or deterioration in a distributed batch is discovered, or when a bacterial contamination or a mislabelling is found — within three working days of becoming aware. Three working days is short enough that the process must be ready in advance, which is exactly the kind of requirement that turns into a workflow system.
Escalation: who gets told, and when
Escalation is the formal path by which a problem reaches people with authority to act, and it is written down rather than left to judgement.
A typical structure has three tiers. Routine events are handled by the area's own quality staff. Significant events — a critical deviation, a confirmed out-of-specification result on released product, a repeated failure — go to site quality leadership within a defined time. And the highest tier goes to corporate quality and executive management: anything suggesting product in the market may be unsafe, anything that could become a recall, a regulatory inspection finding, or a data integrity concern.
Two things make escalation work or fail, and both are cultural rather than procedural.
The trigger must be defined so that escalating is not a judgement call the person making it can be punished for. If escalation depends on someone's courage, it will be inconsistent, and the events that most needed it are exactly the ones nobody wanted to raise.
And the response to escalation must be support rather than blame. Every serious quality failure in this industry's history was known to someone before it became public. The distance between "known to someone" and "known to the person who could act" is the thing escalation exists to close.
The people, and the roles you will be introduced to
Quality departments have their own job titles and the acronyms fly quickly. Here are the ones you will actually meet.
Quality Assurance owns the system: procedures, approvals, release, audits, CAPA oversight. Quality Control runs the laboratory that tests things. The process owner is the person accountable for a given process working correctly. The subject matter expert — the SME — is the person with the deepest working knowledge of a specific process, system or piece of equipment, and they are the person an investigation, a validation or an inspector will ask for. When someone says "we need to book the SME", they mean the one person who genuinely knows how that thing behaves.
And there is usually a coordinator role — titled variously as quality coordinator, process coordinator or programme coordinator, and abbreviated in company-specific ways — whose job is to keep records moving: triaging incoming events, assigning owners, chasing due dates, preparing the review meeting, and reporting metrics. Acronyms for this role vary between companies and even between sites, so the professional move is to ask what a given abbreviation means at that client rather than assume, since guessing wrong in a meeting is worse than asking.
The metrics that get watched
Every quality system reports a standard set of numbers upward, and knowing them tells you what your client is being judged on.
| Metric | Why it is watched |
|---|---|
| Open deviations, and their age | Backlog means loss of control |
| Overdue CAPAs | Strongest predictor of a bad inspection |
| Right-first-time batch rate | Quality of the process itself |
| Complaint rate per units shipped | Product performance in the field |
| Repeat events | Whether causes are really being fixed |
| On-time closure of records | Whether the system is resourced |
Notice what these have in common. They are all trivially computable from data the company already holds, and in a great many companies they are still assembled by hand into slides every month. That gap — data that exists, insight that does not — is the most reliable place for a services company to demonstrate value quickly and without touching a validated process at all.
Next: Chapter 25.18, what happens when somebody comes to check — internal audits, supplier audits, and the regulatory inspection, including the exact ladder of consequences from an observation to a consent decree.