qsteadgroup@yahoo.com (619) 718-1853

Case Study · CAPA & Investigations

Pipette OOS: When Two Things Change at Once

A look back at a real calibration CAPA: what the team did, what the evidence could and couldn’t prove, and how a stronger investigation would have separated competing causes.

This case is drawn from a real CAPA and rebuilt afterward from memory as a learning exercise. Manufacturer and laboratory names are anonymized. No numerical results are shown, because the original values weren’t available and none have been invented. Throughout, remembered facts, hypotheses, and recommendations made in hindsight are kept separate.

At a glance

ProblemRecurring as-found out-of-specification (OOS) results on pipettes returning from scheduled calibration
SettingQC Chemistry, Microbiology, and R&D laboratories
Tools usedChange Analysis, Kepner-Tregoe Is / Is-Not
Historical actionsCalibration interval cut from 12 to 6 months; an overly tight specification later revised
ConclusionRoot cause not conclusively established. The actions reduced risk, but didn’t prove why the failures happened.

1. What was observed

For years, the labs used pipettes from Manufacturer A, calibrated by external Lab A. The organization then bought a new population of pipettes from Manufacturer B, calibrated by external Lab B. Before long, individual Manufacturer B pipettes began coming back from scheduled calibration with as-found OOS results.

  • The failures looked intermittent and random.
  • They showed up in Chemistry, Microbiology, and R&D, with no concentration in one user, lab, application, or asset.
  • Manufacturer A pipettes weren’t remembered as having a comparable pattern.

Problem statement

Recurring as-found OOS results appeared in the Manufacturer B population after the organization changed both its pipette manufacturer and its external calibration lab. The evidence doesn’t show which change, or what other variable, caused the pattern.

2. What changed, and what didn’t

ElementBeforeAfterWhat it means
Pipette manufacturerManufacturer AManufacturer BChanged. High-priority candidate.
Pipette populationExisting, older unitsNewly purchased unitsChanged along with the manufacturer, so age effects are mixed in too.
Calibration labLab ALab BChanged. High-priority measurement-system candidate.
Calibration interval12 months12 months (later 6)No change at first.
Internal calibration procedureSameSameUnchanged, but that doesn’t prove Lab A and Lab B tested the same way.
Chemistry and Micro staffSame peopleSame peopleStable, which makes the operator a less likely cause.

“Our procedure stayed the same” describes the internal process only. The two outside labs could still have differed in method, test points, replicates, conditioning, measurement uncertainty, decision rules, and how they reported as-found results.

3. Change Analysis came first

With a clear before-and-after transition, Change Analysis was the most useful first tool. It pointed straight at two changes: the manufacturer and the calibration lab. Users, labs, the internal procedure, and the original interval hadn’t changed, so they ranked lower.

4. Kepner-Tregoe Is / Is-Not came next

Change Analysis asks what changed over time. Is / Is-Not asks what separates the problem from the non-problem right now.

IsIs notSo what?
WhatSome Manufacturer B pipettes failed as-foundNot every unit; others passed“The brand is bad” isn’t supported
WhereChemistry, Micro, R&DNot one labLocation doesn’t discriminate
WhoMany experienced usersNot one operatorOperator is lower priority, not ruled out
WhenFound at scheduled calibrationNo known start dateAnnual checks can’t show when drift began
ExtentMultiple assetsNo cluster by asset IDMore detailed data needed

Finding no strong distinction is itself a result. It sent the investigation back to the two simultaneous changes, and toward controlled testing instead of more brainstorming.

5. The core problem: confounding

History only offered two of four possible combinations:

Lab ALab B
Manufacturer ALittle or no OOS rememberedNever tested
Manufacturer BNever testedRecurring OOS

With only the diagonal filled in, you can see that something changed with the transition, but you can’t tell whether it was the pipettes, the lab, the combination, or something else. “Manufacturer B caused it” and “Lab B caused it” were both premature. A third possibility: both labs got similar numbers but applied different acceptance criteria, uncertainty rules, or reporting.

6. Why the usual tools weren’t the first choice

  • Fishbone: would list the familiar categories without separating the two confounded changes.
  • 5 Whys: premature. Without a verified first link, it turns a plausible story into an unsupported chain.
  • Fault tree: too elaborate for what was known about the failure mechanism.
  • FMEA: good for future risk and controls, not for finding a past cause.
  • Pareto and Shainin Red X®: useful later, once failures are categorized or consistent good and bad units are available.

The right order was Change Analysis and Is / Is-Not first, then stratifying and trending the data, then a controlled comparison to test the hypothesis, and only then 5 Whys to work upstream from a demonstrated mechanism.

7. The experiment that would break the confounding

Recommendation made in hindsight, not an action taken at the time

When Lab B reports an as-found OOS, stop. Keep the pipette exactly as it is, with no adjustment, repair, or recalibration. Then send it to a qualified second lab for blinded as-found testing.

If…Then…
Both labs get similar OOS valuesLab B alone is less likely. Look harder at the pipette or at criteria both labs share.
Lab B says OOS, the second lab gets clearly in-spec valuesLook at the measurement system: method, transport, conditioning, the lab itself.
Numbers agree but pass/fail calls differThe issue is acceptance criteria, uncertainty, guard bands, or decision rules.
Results sit right at the limitA borderline system where uncertainty dominates. Neither lab is necessarily wrong.

To make the comparison credible: compare raw values, uncertainty, test points, replicates, and conditions, not just PASS/OOS labels. Blind the second lab to the first result. Define transport, equilibration, and timing ahead of time. And use the same physical pipette as its own comparator. Retired Manufacturer A units aren’t needed: the current population can isolate the lab variable on its own.

8. The historical actions, credited fairly

Cutting the interval from 12 to 6 months

A sensible interim control. It finds problems sooner, reduces how long a drifting pipette stays in use, and creates more data for trending. What it can’t do is show when the problem started, prevent a pipette from going OOS, or identify the cause. It’s best classified as containment or risk control, unless evidence shows the interval itself was the system gap.

Revising an overly tight specification

A scientifically justified revision, based on intended use, method needs, measurement uncertainty, and manufacturer performance, can be the right call. But it moves the pass/fail line without changing the pipette or the measurement.

The key distinction

Fewer OOS results after a spec change show an effect on classification. They don’t explain why the original values occurred. “Root cause: specification too tight” only holds if the problem you defined was wrong classification, not pipette performance.

9. The missing piece: historical Manufacturer A data

The biggest open question is whether the old and new populations were ever judged on comparable terms. These records would change the picture:

  • The acceptance criteria in effect during the Lab A years
  • Lab A certificates with raw as-found results, and whether as-found was always reported
  • Lab B criteria, certificates, test points, raw data, uncertainty, and decision rules

If the old criteria were looser, part of the “new” problem may be a classification effect. If they were the same, “spec too tight” gets weaker as an explanation. If the records can’t be found, the historical cause stays uncertain, and current controls can still be justified on risk.

10. A stronger investigation, step by step

  1. Define the problem: the actual as-found failure and the affected population.
  2. Change Analysis: identify the manufacturer and lab changes.
  3. Is / Is-Not and stratification: compare failing and passing units; trend the raw history.
  4. Check comparability: methods, limits, uncertainty, and decision rules at both labs.
  5. Controlled confirmation: preserve as-found condition; paired interlaboratory testing.
  6. Mechanism, then upstream cause: only now apply 5 Whys or targeted tools.

Supporting evidence: calibration data by asset (model, dates, time in service, as-found and as-left values, repair findings); repeat-failure patterns; physical findings such as seals, pistons, leaks, or damage; before-and-after data for both actions; and a written statement of residual uncertainty if no single cause is confirmed.

11. Verdict

Done well

  • Responded to recurring OOS instead of ignoring it
  • Shorter interval plausibly reduced risk
  • A justified spec revision can be sound science

Unresolved

  • No confirmed physical or measurement mechanism
  • Manufacturer and lab stayed confounded
  • Comparable historical data were missing

What would help

  • Recover comparable historical records
  • Trend current data and repair findings
  • Preserve as-found units for independent testing

Lessons for any investigation

  • Define the problem precisely. A classification problem and a performance problem are different problems.
  • When two factors change together, history alone can’t separate them.
  • Use Change Analysis to find change points, Is / Is-Not to find distinctions, and controlled testing when observation can’t break the tie.
  • Preserve as-found evidence before anyone repairs or adjusts it.
  • Treat an interval cut as risk control unless it’s shown to fix a causal gap.
  • A spec revision can be right even when the physical root cause stays open.
  • “Root cause not conclusively established” is more defensible than a single unsupported explanation.

Related reading: Advanced Root Cause Analysis and Five signs your CAPA system is producing paperwork, not improvement.

Kepner-Tregoe and Shainin/Red X® are proprietary methods of their respective owners, described here for educational purposes only. QStead Group is not affiliated with them. This case study is a retrospective learning example, not a substitute for the original CAPA record.

Working through a calibration, OOS, or CAPA investigation that won’t close? Get in touch.

← Back to Case Studies