qsteadgroup@yahoo.com (619) 718-1853

Education

Advanced Root Cause Analysis

Why investigations stall at “operator error,” and the tools seasoned investigators use to find what really went wrong.

Root cause analysis (RCA) is the hardest part of CAPA (corrective and preventive action). It asks people to look past the obvious answer, and sometimes past their own process. It is also where CAPA most often falls short: FDA inspection findings on CAPA regularly cite investigations that stop at “operator error” or “retrained,” and problems that keep coming back.

This guide covers where RCA usually fails, how to get it right, and the advanced tools that go beyond 5 Whys and the fishbone diagram.

Where root cause analysis usually fails

  • Stopping at the first answer. “The operator skipped a step” is what happened, not why. Why was skipping possible? Why wasn’t it caught?
  • Blaming people instead of systems. Human error is almost always a symptom. Unclear instructions, poor equipment design, time pressure, and weak training content make mistakes easy.
  • Deciding the answer first. Under deadline pressure, teams pick the cause that is quickest to fix, then work backward to justify it.
  • Investigating alone. One investigator at a desk misses what operators, engineers, and suppliers already know.
  • Opinions instead of data. Without batch records, trend data, environmental data, or a timeline, RCA becomes guesswork.
  • One-and-done thinking. Each event is closed on its own, so nobody notices the same cause behind five different deviations.

How to get it right

  1. Define the problem precisely. What, where, when, how big, and what didn’t happen. A vague problem statement leads to a vague root cause.
  2. Use a structured tool. Tools make you consider causes you would otherwise skip.
  3. Go to the floor. Walk the process, talk to the people who do the work, and look at the actual equipment and records.
  4. Test the root cause. Ask: “If we fix this, can the problem still happen?” If yes, keep digging.
  5. Separate the fix from the cause. Retraining can be part of a fix. It is rarely the root cause.
  6. Trend across events. Review deviations, complaints, nonconformances, and audit findings together to find causes that keep coming back.
  7. Prove it worked. Set measurable effectiveness criteria before closing the CAPA.

The advanced toolkit

5 Whys and the fishbone (Ishikawa) diagram are good starting points for simple problems. For complex, recurring, or high-risk problems, experienced investigators reach for tools like these, grouped by the job each one does.

1. Framing the problem precisely

  • Kepner-Tregoe Problem Analysis (Is / Is-Not). Compares where, when, and how much the problem is happening with where it logically could happen but isn’t. The differences point directly at the cause. Best for: sporadic or confusing failures.
  • Change Analysis. Compares a good period or batch with a bad one and lists everything that changed: materials, people, equipment, methods, environment. Best for: a process that was working and suddenly isn’t. Most failures follow a change.

2. Mapping how the event happened

  • Events & Causal Factors (E&CF) charting. A timeline of events and the conditions around them. Best for: complex deviations and serious incidents.
  • Cause Mapping / Apollo-style RCA. A cause-and-effect map in which every cause must be supported by evidence. Unlike a fishbone, it shows how causes combine. Best for: problems with several contributing causes.
  • Fault Tree Analysis (FTA). Works top-down from the failure using AND/OR logic, and can be made quantitative. Best for: device and equipment failures.
  • Bowtie Analysis. Shows the causes, the controls meant to stop them, and the consequences on either side of a single event. Best for: connecting investigations to your ISO 14971 risk management file.

3. Finding which controls failed

  • Barrier Analysis. Asks which controls (SOPs, interlocks, inspections, reviews) should have stopped the problem, and why each one didn’t. Best for: “How did this get past QC?”
  • HFACS (Human Factors Analysis and Classification System). A structured way to analyze human error at four levels: unsafe acts, preconditions, supervision, and organization. Best for: replacing “operator error” with a real answer.
  • TapRooT®. A commercial RCA system built around a root cause tree and human performance. Best for: organizations that want one standardized, trained method.

4. Letting the data decide

  • Shainin methods (Red X®). Multi-vari charts, component search, and paired comparisons find the dominant cause by comparing best and worst parts, without large studies. Best for: chronic manufacturing variation.
  • Design of Experiments (DOE). Confirms the cause by deliberately switching the problem on and off. Best for: proving a root cause, the gold standard.
  • SPC, Pareto, regression, and hypothesis testing. Separate signal from noise and show whether a suspected cause really correlates with the failure. Best for: any investigation with data behind it.
  • Weibull analysis. Shows whether failures are early-life, random, or wear-out. Best for: device reliability, field returns, and equipment.

5. Physical evidence

  • Failure analysis lab work. Microscopy, SEM/EDS, FTIR, DSC, fractography, cross-sectioning, and CT scanning show physically what failed and how. Best for: material, component, and device failures, where evidence beats opinion.

6. Frameworks that tie the tools together

  • 8D (Eight Disciplines). Team-based problem solving with containment, root cause, correction, and prevention steps. Often required by customers and OEMs.
  • DMAIC (Six Sigma). Define, Measure, Analyze, Improve, Control. Suited to chronic, data-rich problems.
  • Reverse / process FMEA. Feeds what the investigation found back into the risk file so the lesson carries forward.

Match the tool to the problem

  • Sporadic or intermittent defect: Is / Is-Not, Change Analysis
  • Equipment or device failure: Fault Tree Analysis, Weibull, failure analysis lab work
  • Human-performance event: HFACS, Barrier Analysis
  • Chronic process variation: Shainin methods, SPC, DOE
  • Serious or multi-cause incident: E&CF charting, Cause Mapping, Bowtie
  • Customer-driven complaint: 8D

What sets seasoned investigators apart

They don’t use one tool for everything. They choose based on the type of problem, and they verify the cause before closing the CAPA, either by reproducing the problem or by showing that the fix eliminated it. Good RCA isn’t about finding someone to blame. It’s about finding what in the system allowed the problem to happen.

Related reading: Five signs your CAPA system is producing paperwork, not improvement and QMS for Beginners.

Kepner-Tregoe, Apollo Root Cause Analysis, TapRooT®, and Shainin/Red X® are proprietary methods of their respective owners. They are described here for educational purposes only. QStead Group is not affiliated with them.

Stuck on an investigation, or want a second set of eyes on your CAPA system? Get in touch.

← Back to Education