Cleaning validation is the documented evidence that your cleaning procedures remove drug product residues, and the residues of the cleaning agents themselves, from manufacturing equipment to a level that does not pose a risk to the next product or to a patient. It applies to any multi-product equipment, meaning any piece of equipment used to make more than one product and cleaned between campaigns. It also applies, in a narrower way, to dedicated equipment, where the question shifts from “what carries over to the next product” to “does residue or degradant build-up affect the same product over time.”
The same framework runs across modalities. A small-molecule tablet line, a biologic bioreactor train, a medical-device assembly that contacts product, and a sterile fill suite all share the carryover problem. The numbers and the dominant risk change (chemical carryover for small molecules, protein and endotoxin carryover for biologics, particulate and bioburden for devices), but the logic of limit, sample, recover, and trend is identical.
The regulatory basis rests on a small set of documents. In the US, 21 CFR 211.67 (equipment cleaning and maintenance) requires written cleaning procedures and records. The expectation that those procedures be validated comes from FDA’s 1993 Guide to Inspections, Validation of Cleaning Processes. In Europe, EudraLex Volume 4, Chapter 3 and Chapter 5 require that cross-contamination be prevented by adequate cleaning, and the 2014 EMA guideline on setting health-based exposure limits for use in risk identification in the manufacture of different medicinal products in shared facilities (EMA/CHMP/CVMP/SWP/169430/2012) made health-based limits the expected approach. That guideline replaced the older visual-inspection-plus-10-ppm habits that had been industry practice for decades. ICH Q7 (Good Manufacturing Practice for Active Pharmaceutical Ingredients) sets parallel expectations for API manufacturing. For medical devices, the quality-system rule (21 CFR 820, transitioning to the harmonized QMSR aligned with ISO 13485) carries the same cleaning and contamination-control obligations under process validation. If you want the broader picture of how cleaning validation sits inside the validation program, the process validation lifecycle and the validation master plan articles give the surrounding structure.
Why Cleaning Validation Matters
Consider a concrete scenario. Product A is an immunosuppressant with a narrow therapeutic index. Product B is a low-potency oral product with no significant toxicological concern. They share a granulation line. The equipment is cleaned between campaigns, and Product B is the next product to run.
If the cleaning is inadequate and Product A residue carries into Product B, patients taking Product B receive an unintended dose of an immunosuppressant. That is the patient-safety case, and it is not academic. Cross-contamination findings have driven recalls, consent decrees, and the shutdown of multi-product lines. The cost of getting this wrong is not a paperwork observation, it is product on the market that exposes patients to a drug they were never prescribed.
The reverse direction matters too. Cleaning agent residues, the solvents, detergents, and acids or bases used to clean, can themselves contaminate the next batch. A caustic cleaner left on a surface can degrade an acid-sensitive API. So cleaning validation has to address two carryover paths: previous product into next product, and cleaning agent into next product. A third concern, microbial and endotoxin carryover, becomes the dominant one in sterile and biologic operations and links directly to the aseptic processing and media fills and environmental monitoring programs.
Three carryover paths to keep straight, because the limit and the test method differ for each:
| Carryover path | What you are limiting | Typical test |
|---|---|---|
| Previous product into next product | API or formulation residue | Specific assay (HPLC), or TOC as a non-specific surrogate |
| Cleaning agent into next product | Detergent, acid, base, solvent | TOC, conductivity, or a component-specific method |
| Microbial / endotoxin carryover | Bioburden and bacterial endotoxin | Plate count, bioburden and endotoxin testing |
Acceptable Daily Exposure (ADE) and Permitted Daily Exposure (PDE)
ADE (the term ISPE and FDA practice tend to use) and PDE (the EMA term) are numerically the same concept. Both represent the maximum amount of a substance a patient can be exposed to per day without an appreciable risk of adverse effects, derived from toxicological and pharmacological data. The two terms are interchangeable in practice, and a single value drives the limit on either side of the Atlantic. ICH Q3C (residual solvents) and ICH Q3D (elemental impurities) use PDE the same way for those impurity classes, so the calculation style is consistent across the impurity guidances.
Why this matters for cleaning: if you know the ADE of the previous product, you can back-calculate the maximum residue that can sit on the cleaned equipment and still keep a patient’s exposure through the next product below that ADE.
The ADE is derived as:
ADE = NOAEL × BW / (UF₁ × UF₂ × UF₃ × UF₄ × UF₅)
Where:
- NOAEL is the No-Observed-Adverse-Effect Level from the most relevant study, expressed in mg/kg/day. When no clean NOAEL exists, a LOAEL (Lowest-Observed-Adverse-Effect Level) is used with an extra adjustment factor, or a clinical therapeutic dose is used as the point of departure.
- BW is body weight, defaulted to 50 kg for a human (a conservative choice from the EMA guideline rather than the older 70 kg adult)
- UF₁ through UF₅ are uncertainty factors applied as divisors: interspecies extrapolation (commonly 2 to 12), interindividual variability (commonly 10), study-duration adjustment (subchronic to chronic, up to 10), severity and reversibility of the effect, and a variable factor for an incomplete point of departure, applied when a NOAEL was not established and a LOAEL had to be used instead. This last factor follows the ICH Q3C convention. If you want a separate allowance for an overall thin or poor-quality database, treat that as an additional modifying factor rather than relabeling F5.
The result is in mg/day. For genotoxic or carcinogenic compounds with no identifiable threshold, you do not use a NOAEL at all. You either apply the Threshold of Toxicological Concern (commonly 1.5 µg/day for a non-cohort-of-concern genotoxicant) or extrapolate linearly from carcinogenicity data to a 1-in-100,000 risk level. ICH M7, on mutagenic impurities, is the reference here, and the same logic appears in the nitrosamines and mutagenic impurities work. Highly sensitizing compounds (some beta-lactams, certain cytotoxics) are handled separately again: regulators expect dedicated, not just well-cleaned, facilities for those, which the cross-contamination control in shared facilities article covers.
A worked example makes the chain concrete. Suppose the previous product has a NOAEL of 1 mg/kg/day from a 90-day rat study, and you apply factors of 5 (interspecies), 10 (interindividual), 5 (subchronic to chronic), 1 (effect not severe), and 1 (a clean NOAEL, so no LOAEL adjustment). The composite factor is 250.
ADE = (1 mg/kg/day × 50 kg) / 250 = 0.2 mg/day = 200 µg/day
The single most important caution: ADE derivation is a toxicology task, not a quality task. It must be done by a qualified toxicologist, documented in a signed monograph, and the basis for each factor must be visible. Pulling a generic “10 ppm” limit out of the air with no toxicological reasoning is not defensible under current expectations.
What goes in an ADE monograph. When an inspector asks for the basis of a limit, this is the document they want to see. A defensible monograph contains: the compound identity and a short pharmacology and toxicology summary; the critical study and the chosen point of departure (NOAEL, LOAEL, or clinical dose) with the citation; each uncertainty factor and a sentence justifying its value; the resulting ADE in mg/day; a statement on genotoxic, sensitizing, or reproductive-hazard potential; the author’s qualifications; and a signature and date with a review and re-issue interval. An ADE without traceable reasoning behind each factor is treated as no ADE at all. The HBEL/PDE monograph template builds this document field by field, with a worked derivation you can adapt.
From ADE to Equipment Residue Limit
Once you have the ADE in mg/day, calculate the Maximum Allowable Carryover (MACO), the total mass of previous product that may carry into a full batch of the next product:
MACO (mg) = ADE_previous (mg/day) × Batch size of next product / Maximum daily dose of next product
The units have to line up: batch size in the same mass unit as the dose. If the next product has a 200 kg batch and a maximum daily dose of 400 mg, and ADE_previous is 0.2 mg/day:
MACO = 0.2 mg/day × (200,000,000 mg / 400 mg/day) = 100,000 mg = 100 g
That MACO is then spread over the shared equipment surface to give a surface limit:
Limit (µg/cm²) = MACO (µg) / Total shared product-contact surface area (cm²)
If the shared train has 500,000 cm² of product contact area:
Limit = 100,000,000 µg / 500,000 cm² = 200 µg/cm²
That number is the per-area target the swab result has to beat. Health-based limits sometimes come out higher than the old defaults, sometimes much lower. The point of the ADE approach is that the limit is now tied to the actual toxicity of the molecule rather than an arbitrary blanket factor. The three historical methods are worth knowing because you will still see them and may use the most conservative as a cross-check.
| Method | Basis | Status |
|---|---|---|
| Health-based (ADE/PDE) | Toxicological NOAEL or TTC | Current expectation (EMA 2014, ICH Q7) |
| 0.1% of therapeutic dose | 1/1000 of the previous product’s lowest daily dose carried into the next | Legacy; may be used as an additional cross-check |
| 10 ppm | 10 mg of previous product per kg of next product | Legacy; not acceptable as the sole basis |
Common practice is to calculate the limit by the ADE method, optionally calculate the legacy methods, and adopt the most stringent. That defends the limit against either a regulator who wants health-based reasoning or one who remembers the old conventions.
Converting the surface limit into a swab acceptance number. The surface limit in µg/cm² is not what the lab reports. The lab reports a concentration in the swab extract. You have to carry the limit through the swab area, the extraction volume, and the recovery factor so the analyst knows the pass/fail value in the units the instrument produces. Worked end to end, using the 200 µg/cm² surface limit, a 25 cm² swab area, a 10 mL extraction volume, and an 80% swab recovery:
- Mass allowed on the swabbed area = 200 µg/cm² × 25 cm² = 5000 µg
- Corrected for 80% recovery = 5000 µg × 0.80 = 4000 µg actually recovered into the vial at the limit
- Concentration in the extract = 4000 µg / 10 mL = 400 µg/mL
So a swab result above 400 µg/mL fails. Note the limit above is unusually high because the example molecule is benign; a potent compound can drive the swab acceptance value down to single-digit µg/mL or into ng territory, which is exactly where method sensitivity becomes the constraint. Always pre-calculate this number, put it in the protocol, and confirm the analytical method’s limit of quantitation sits comfortably below it before you run a single swab.
The whole chain, walked above in numbers, reduces to five steps every time:
Work the calculation in that order every time. The MACO and acceptance limit calculation worksheet is built to walk exactly this sequence with your own numbers, cross-checked against the 10 ppm and dose criteria before you commit to a limit.
Dedicated vs Shared Equipment: The Upstream Decision
Before worst-case selection even starts, there is a prior question: should this product be on shared equipment at all? Cleaning validation proves a shared line is safe to share. It does not answer whether sharing was the right call in the first place, and that decision sits in quality risk management. EU GMP Chapters 3 and 5 (2015 revision) expect the dedication-versus-shared decision to be documented, not assumed, and it is one of the first things an inspector probes when a facility runs a mixed portfolio.
Why some products get dedicated equipment. Certain beta-lactam antibiotics, some sex hormones and cytotoxics, and other highly sensitizing or potent compounds carry a long regulatory history of dedicated-facility expectations, because sensitization and potency make even a trace, validated-limit-compliant carryover a real patient-safety event rather than a theoretical one. For these, no amount of cleaning validation rigor substitutes for physical separation.
What drives the decision. Four inputs matter: the ADE/PDE and its genotoxic or sensitizing classification; the acceptance limit the ADE chain above would produce on the intended shared train; whether the existing, or realistically achievable, cleaning procedure has ever been demonstrated to reach that limit with margin, not just assumed to; and whether the analytical method’s limit of quantitation actually sits below that limit. If the honest answer to either of the last two is no, tightening the cleaning SOP on paper does not fix it. Dedication does.
How to work the decision, step by step.
- Derive the ADE/PDE and the genotoxic or sensitizing classification for the candidate product from the signed toxicology monograph.
- Run the ADE-to-swab-limit chain from the section above for the intended shared train to see what the acceptance number would actually be.
- Compare that number to the best demonstrated recovery-corrected cleaning performance on the train, and to the analytical method’s proven limit of quantitation in that matrix.
- Weigh the hazard classification on its own terms. A compound can fail the sharing test on sensitization or genotoxicity grounds even when the arithmetic looks achievable.
- Document the decision, including the alternative that was rejected and why, in a quality risk management record, and route it to Quality for approval before the product is scheduled.
- Re-trigger the assessment whenever a new product is proposed for the train, new toxicology data changes the ADE, or cleaning performance trends show the margin eroding.
Acceptance criteria for the decision itself: the quality risk management record names the product; states the hazard classification and its source; shows the calculated acceptance limit against the demonstrated cleaning and analytical capability; states the decision and the rationale for it, including alternatives considered; carries Quality approval before manufacture; and defines the triggers that force reassessment.
A worked example. A potent hormone-class compound has an ADE of 0.4 µg/day, driven by a low point of departure and a large composite uncertainty factor. On the intended shared solid-dose train (shared surface area 300,000 cm², swab area 25 cm², recovery 75 percent), the chain above gives: MACO = 0.4 µg/day × 150,000,000 mg (minimum next batch) / 200 mg/day (largest next dose) = 300,000 µg = 0.3 g; surface limit = 300,000 µg / 300,000 cm² = 1.0 µg/cm²; swab-area mass = 1.0 µg/cm² × 25 cm² = 25 µg; recovery-corrected acceptance = 25 µg / 0.75 = 33.3 µg per swab, roughly 3.3 µg/mL in a 10 mL extract. The site’s validated HPLC method for this compound in this matrix has a demonstrated limit of quantitation of 8 µg/mL, more than double what the limit requires. A “not detected” result from that method would prove nothing. Rather than accept it, the site dedicates a smaller train to the product and documents the limit-of-quantitation gap as the deciding factor.
Roles. Toxicology provides the hazard classification and the ADE. Validation and QC analytical jointly demonstrate, rather than assert, achievable cleaning performance and method sensitivity. Quality risk management owns the decision framework and Quality Assurance approves the outcome. Manufacturing and facilities carry the real cost and schedule consequence of a dedication decision, and that pressure should never be allowed to move the technical answer.
For the facility-design and containment controls that sit alongside this decision, see cross-contamination control in shared facilities.
Worst-Case Product Selection
In a multi-product facility you cannot validate every product sequence. The worst-case approach picks the single most demanding combination and argues that everything else is covered by it.
- Worst-case previous product (the contaminant): the product with the lowest ADE, hardest solubility, and highest toxicity that runs on the equipment. The lowest ADE sets the most stringent limit. Solubility and stickiness drive how hard it is to remove. These two can point at different products, so document both and justify the choice.
- Worst-case next product (the receiver): the product with the smallest batch size and the highest maximum daily dose. A small batch concentrates the carryover, and a high daily dose means more of the contaminated product reaches the patient per day.
- Worst-case equipment: the unit in the train that is hardest to clean. Complex geometry, the most product-contact area, the smallest internal clearances, dead legs, long transfer lines, and porous or scratched surfaces all push an item toward worst case.
A clean way to organize this is a risk-ranked matrix that scores each product on hardest-to-clean (solubility), toxicity (ADE), and dose/batch attributes. Quality risk management methodology, covered in quality risk management, gives the scoring discipline so the worst-case choice is reproducible rather than a matter of opinion. Validation run against the worst case is taken to represent the worst cleaning challenge, and other sequences are addressed by documenting that they are less demanding.
A worked worst-case matrix. Here is the kind of table that belongs in a cleaning validation master plan. Each product on the line is scored, and the winners on the two axes that matter are flagged. A simple convention: rank ADE from lowest (most stringent, highest concern) and solubility from least soluble (hardest to clean).
| Product | ADE (µg/day) | Solubility | Max daily dose (mg) | Batch size (kg) | Cleanability rank | Toxicity rank |
|---|---|---|---|---|---|---|
| Alpha | 50 | Poor | 200 | 150 | Hardest (1) | High (2) |
| Beta | 20 | Moderate | 100 | 80 | Medium (2) | Highest (1) |
| Gamma | 500 | Good | 400 | 250 | Easiest (4) | Low (4) |
| Delta | 120 | Poor | 50 | 60 | Hard (2) | Medium (3) |
Reading this matrix: Beta drives the acceptance limit (lowest ADE, so the tightest µg/cm²), while Alpha drives the cleaning challenge (poorest solubility, hardest to physically remove). Best practice is to validate the cleaning procedure against the harder-to-clean soil (Alpha) but apply the more stringent limit (from Beta) to the result, so the worst case on both axes is covered. The worst-case next product needs more care than the slogan “smallest batch, highest dose” suggests, because those two attributes can point at different products. Since MACO scales with batch size and inversely with daily dose, the product that squeezes MACO hardest is the one with the smallest batch-to-dose ratio. Compute that ratio for each: Alpha 150/200 = 0.75, Beta 80/100 = 0.80, Gamma 250/400 = 0.625, Delta 60/50 = 1.2. Delta has the smallest batch but its low 50 mg dose works the other way, so its ratio is the largest and it gives the most forgiving MACO. Gamma, despite a mid-size batch, has the smallest ratio and is the worst-case next product. When smallest-batch and highest-dose flag different products, calculate the ratio and pick the smallest, do not default to the smallest batch alone. Document each of these three choices with the reasoning, not just the conclusion.
Turn this scoring exercise into a living document, not a one-time table: the worst-case grouping matrix template carries the product property table, an A-into-B MACO transition matrix for every credible pairing, and the re-trigger logic that recomputes the governing case the moment a new product joins the train.
Cleaning Validation Protocol Design
Worst-case soil. Validation runs at the maximum soil load the equipment will see in routine production: the largest batch, the highest product loading, and the longest dirty-hold time allowed between end of processing and start of cleaning. Soil that has dried on a surface for the maximum permitted dirty-hold is harder to remove than fresh soil, so the protocol has to fix and challenge that hold.
Clean-hold and dirty-hold times. Two time windows belong in the protocol. The dirty-hold time (DHT) is how long equipment may sit soiled before cleaning starts. The clean-hold time (CHT) is how long cleaned equipment may sit before it must be used or re-cleaned, and the CHT is usually bounded by microbial growth, established with bioburden and endotoxin testing rather than chemical residue. Both windows are validated at their maximum: the protocol holds soiled equipment for the longest DHT before cleaning, and holds cleaned equipment for the longest CHT before sampling, so the limits printed in the SOP are the ones actually proven.
The procedure is fixed first. Define the exact cleaning procedure before validation: agent, concentration, temperature, contact time, flow rate or action, rinse volume, and rinse quality. Validation confirms that the written procedure works. If validation needed five rinses but the SOP says three, the SOP, not the better run, is what production will follow, and the validation did not represent reality. Any change to the procedure afterward goes through change control with an assessment of whether revalidation is needed.
Protocol contents. A complete cleaning validation protocol contains, at minimum: purpose and scope (which equipment train, which products); references to the cleaning SOP being validated and the ADE monographs; the worst-case justification for product and equipment; the calculated MACO, surface limit, and swab acceptance concentration with the math shown; the sampling plan with swab location maps and rinse points; the analytical methods and their validation status; the recovery study reference; the number of runs (conventionally three) and the conditions for each; acceptance criteria for chemical residue, cleaning agent, bioburden, endotoxin, and visual; roles and signatures; and the deviation-handling approach. Approve the protocol before execution. Generating data and then writing the protocol to match is a serious data-integrity problem, not a shortcut. The cleaning validation protocol template assembles all of this into one ready-to-use document, with the calculation, sampling, and acceptance sections already structured.
The validation summary report. Execution produces the report, and it is the document an inspector reads most closely, because it states the conclusion. A complete report presents the actual recovery-corrected result at every sampled location, not just pass or fail; lists every deviation encountered during execution and its impact on the conclusion, never quietly absorbing one; confirms the DHT and CHT actually challenged, as opposed to what the protocol merely hoped to demonstrate; and states plainly whether the cleaning process is validated, partially validated with follow-up commitments, or not validated. Writing “all results passed” with no numbers shown is a common and avoidable gap that undermines an otherwise sound study. The cleaning validation summary report template gives the section-by-section structure this needs, including the bridge into routine verification once the study closes.
Sampling strategy. Two methods, usually together.
Swab sampling wipes a swab, pre-moistened with an appropriate solvent, over a defined area (commonly 25 cm²) on product-contact surfaces, then extracts and analyzes it. Swabs measure residue at specific, deliberately chosen locations. Location selection has to be worst case: gaskets, welds, baffles, dead legs, agitator shafts, valve seats, and any geometry that traps soil. The swab locations are mapped and the map goes in the report so an inspector can see exactly where you sampled. A consistent swabbing technique matters as much as the location: a defined number of strokes, in two directions, with a fixed pressure, trained and documented, because recovery is technique-dependent.
Rinse sampling collects and analyzes the final rinse. Rinse covers the whole wetted surface but dilutes the residue and can miss a localized hot spot in a hard-to-reach pocket. Rinse is good for surfaces a swab cannot reach, such as the inside of long transfer lines, but rinse alone is generally not enough. The two methods complement each other: swabs for direct, localized measurement, rinse for whole-surface coverage. The swab and rinse sampling plan template sets out the numbered-location convention, the technique detail, and the blanks and chain-of-custody controls that make sampling reproducible across samplers and runs.
Analytical methods. The method has to be validated for that analyte in that matrix. A validated HPLC assay for the API in finished product is not automatically valid for swab extracts off stainless steel. Specific methods (HPLC, LC-MS for trace work) quantify a named residue; non-specific methods such as total organic carbon (TOC) and conductivity measure everything and are useful for rinse and for cleaning agents. The method’s limit of quantitation must sit comfortably below the residue limit you calculated. Method validation principles are covered in method validation essentials.
The table below lines up the detection options side by side, because picking the wrong one for the residue in front of you is a recurring, avoidable error.
| Method | What it measures | Specificity | Typical sensitivity | Best used for |
|---|---|---|---|---|
| Visual inspection | Presence of any visible residue under defined conditions | None (qualitative) | Roughly 1 to 4 µg/cm² for many APIs, worse for potent compounds | First-pass qualitative gate, always required, never the sole quantitative criterion |
| Swab, specific assay (HPLC, LC-MS) | A named compound at a defined location | High | Method-dependent; must sit below the per-swab limit | Worst-case, hard-to-clean locations; potent or genotoxic actives |
| Swab or rinse, TOC | All organic carbon, non-specific | Low | Sub-ppm, very sensitive | Fast screening across multi-product trains; cleaning-agent and degradant carbon |
| Rinse, conductivity or pH | Ionic species, cleaning-agent rinse-off | Low, ionic only | Adequate for ionic cleaners; blind to non-ionic surfactants | Confirming cleaning-agent rinse-off, cheap and fast |
| Protein-specific immunoassay (ELISA/HCP) | A named host-cell protein or protein class | High | Low ppm to sub-ppm relative to product, assay-dependent | Biologics on shared protein-contact equipment |
| qPCR (residual DNA) | Host-cell or vector DNA | High | Picogram range | Biologics and gene-therapy vector processes |
No single row is sufficient on its own; a defensible program layers visual with at least one quantitative method, and adds a second method where the first cannot see everything relevant (ionic conductivity misses non-ionic surfactant, TOC does not identify which compound is present).
Swab recovery. How much of what is actually on the surface does the swab-plus-extraction step recover? Spike known amounts onto coupons of the real equipment material (for example 316L stainless steel, but also any elastomer or plastic in contact) at several levels, then swab, extract, and analyze. Recovery should generally be at or above 70% and be consistent across the range. The limit has to be corrected for it: if recovery is 80%, multiply the reported result by 100/80 to estimate the true surface concentration. A recovery below 50% is usually rejected outright; a recovery between 50% and 70% may be accepted with a documented justification and the correction applied, but it weakens the result. Run recovery on every surface material the residue contacts, not just stainless, because recovery off a scratched gasket or a PTFE seal is often much lower than off polished steel. Skipping recovery makes every clean result look better than it is.
Visual inspection. Cleaned equipment should show no visible residue under defined lighting, distance, and viewing angle. Visual is the first gate and a strong tool, but it is not the primary quantitative limit. Visible-residue detection thresholds run roughly 1 to 4 µg/cm² for many APIs, which is fine for low-potency products and not nearly stringent enough for potent ones. For a high-potency compound the calculated limit can be well below what any operator can see, so visual passes while the surface is still over the limit. Define the visual criteria objectively: illumination (lux), distance, angle, the surfaces inspected, who is qualified to perform it, and a documented training that includes spiked-coupon examples of “just visible.”
Roles and Responsibilities
Cleaning validation is cross-functional, and a frequent inspection finding is that nobody could say clearly who owned what. The split below is typical.
| Activity | Owner | Supporting |
|---|---|---|
| ADE/PDE derivation and monograph | Toxicologist | Regulatory, QA |
| Worst-case selection and matrix | Validation / Quality Engineering | QA, Manufacturing |
| Protocol authoring and approval | Validation author; QA approves | Manufacturing, QC, Toxicology |
| Cleaning SOP and execution | Manufacturing | Engineering |
| Sampling (swab, rinse) | Trained sampler (QC or validation) | QA witness for key runs |
| Analytical testing and recovery | QC laboratory | Method validation group |
| Final disposition and report approval | Quality Assurance | All authors |
| Ongoing verification and trending | Manufacturing and QA | Validation |
The principle is that the people who write the limit (toxicology and validation) are independent of the people who run the cleaning (manufacturing), and Quality holds the approval and disposition authority. The GxP roles and responsibilities article covers how these accountabilities are documented in a RACI.
Number of Validation Runs
The 1993 FDA guide established three consecutive successful runs as the conventional minimum. Three passing runs show the procedure is reproducible. A single passing run shows only that it can pass once. All three must run under worst-case conditions throughout: worst-case soil, the maximum dirty-hold, and the exact procedure as written.
The same lifecycle thinking that reshaped process validation is now applied to cleaning. Rather than treating “three runs and done” as the end, mature programs frame cleaning as design, qualification, and ongoing verification, which is why continued monitoring after the three runs has become an expectation rather than a nice-to-have. This mirrors the three-stage model in the process validation lifecycle: cleaning process design, cleaning process qualification (the three runs), and continued cleaning verification.
A subtle point that trips people up: “three consecutive successful runs” means three in a row with no failures in between. If run two fails, you do not get to keep runs one and three and call it a pass. You investigate the failure, fix the cause through deviation management, and start the count again. The conventional three is a floor, not a ceiling; a risk-based justification can support more runs for a difficult worst case or a manual cleaning process with high variability.
Cleaning Agent Residue Limits
Cleaning agent residue has to be shown acceptable too. Three routes:
- Toxicological (ADE) approach: derive an ADE for the cleaning agent from its toxicological profile and treat it like any other residue.
- Component-based approach: for a formulated detergent, assess the most toxic component and set the limit on that, since the formulation as a whole may not have a tidy toxicology dossier.
- Non-detect approach: show the agent is below the validated limit of quantitation in rinse samples, appropriate when that LOQ already sits far below any plausible toxicological concern.
For common agents with established safety profiles, sodium lauryl sulfate, citric-acid-based acids, and many commercial alkaline cleaners, the non-detect route is often enough, and conductivity or TOC on the rinse does the measurement cheaply. For a proprietary or novel cleaner, get the toxicological assessment and the safety data sheet, and confirm the supplier can tell you the components. Supplier documentation here ties into supplier and vendor qualification.
A practical tip for rinse-based agent limits: conductivity is a fast, cheap proxy but it only sees ionic species, so a non-ionic surfactant can sit on the surface invisibly while conductivity reads clean. Match the surrogate to the chemistry. TOC catches organic agents that conductivity misses, and a rinse-water-versus-purified-water conductivity comparison only works for ionic cleaners.
Cleaning Validation in Biologics, Sterile Manufacturing, and Single-Use Systems
The opening of this article said the framework runs across modalities: limit, sample, recover, and trend stay the same, but the numbers and the dominant risk change. That single line is worth unpacking, because biologics, cell and gene therapy, and sterile injectable manufacturing are exactly where a small-molecule-shaped cleaning validation program stops fitting.
Why the risk profile is different. A small-molecule API is a defined chemical entity with a stable ADE derived from dose-response toxicology. A biologic residue is usually a large, complex molecule, a protein, a nucleic acid, a viral vector, or a cellular component, where the dominant carryover risks are host-cell protein (HCP), residual host-cell DNA, endotoxin, and, in cell and gene therapy, viral vector or replication-competent-virus carryover, rather than a classic dose-response toxicity limit. A trace of the wrong protein is more likely to trigger an immunogenic reaction in a patient than a classic organ-toxicity event, and that changes both the limit-setting logic and the assay underneath it.
What changes in the assay and limit approach. Total organic carbon and a specific HPLC assay, the workhorses for small molecules, still have a role, particularly for cleaning-agent residue and buffer components, but the safety-relevant residues usually need different tools: a validated HCP immunoassay (commonly ELISA) referenced against the specific host-cell line and process, quantitative PCR for residual DNA, and the Limulus amoebocyte lysate (LAL) test or a recombinant-factor-C alternative for bacterial endotoxin. A generic total-protein assay, playing a role analogous to TOC, can serve as a non-specific screen across a shared protein-contact train, with the specific HCP or DNA assay reserved for the higher-risk, harder-to-defend claim.
Endotoxin carryover has an unusually clean, citable regulatory hook, worth knowing cold: the threshold pyrogenic dose approach sets the endotoxin limit as K divided by M, where K is the threshold pyrogenic dose in endotoxin units (EU) per kg per hour (5 EU/kg/hour for most parenteral routes, a stricter 0.2 EU/kg/hour for intrathecal administration) and M is the maximum dose of drug product per kg of body weight per hour. That per-mg or per-mL limit then carries through the same surface-area and recovery logic already used for chemical residue, just expressed in endotoxin units instead of micrograms.
A worked example. A biologic drug product has a maximum dose of 2.5 mg/kg administered over roughly one hour by intravenous infusion. Using K = 5 EU/kg/hour: EL = K / M = 5 / 2.5 = 2 EU per mg of drug product. That per-mg endotoxin budget then carries through the same MSSR-and-swab-area chain used for chemical residue on a shared chromatography skid or bioreactor, just in EU rather than µg: divide the total endotoxin budget for the shared batch by the shared surface area to get a per-cm² figure, then multiply by the swab area, then correct for the swab’s endotoxin recovery. The arithmetic discipline is identical; only the unit and the assay change. Numbers here are illustrative; the real limit is calculated per product from its actual maximum dose, route, and labeling, and reconciled against USP <85> and the applicable pharmacopeial endotoxin chapter.
Single-use systems change the question entirely for a large share of biologics processes. Bags, tubing, connectors, and increasingly whole unit operations used once and discarded remove the cross-contamination pathway that cleaning validation exists to close, because there is no shared surface to carry residue between batches. That does not mean no validation work is required: single-use components need extractables and leachables assessment (what the plastic contact surface gives up into the product), integrity and compatibility qualification, and a defined changeover procedure to prevent mix-ups, none of which is a cleaning validation problem in the traditional sense. See extractables and leachables for that discipline. Where stainless-steel equipment remains in the train, most often fermenters, bioreactors, and chromatography skids, standard cleaning validation applies with the biologics-specific assays above layered on top of it.
Roles shift accordingly. Process development and analytical development, not just QC, often own the HCP and DNA assay development and the process-specific reference standards those assays depend on. A biologics-specific safety or immunogenicity risk function may replace or supplement classical toxicology for the hazard framing. Validation engineering still owns the protocol, sampling plan, and report structure, unchanged in shape even though the residues and assays underneath it are different.
Cell and gene therapy pushes this further still: living, patient-specific, or viral-vector-based products raise chain-of-identity and cross-contamination questions that a swab-and-rinse program alone does not answer. See ATMP GMP for cell and gene therapy manufacturing and data integrity in gene therapy for that layer.
Bracketing and Matrixing
Bracketing validates the extremes and infers the middle. Validate the worst-case product and, where it strengthens the argument, the easiest, and treat intermediate products as bracketed by them. This works when the products form a genuine continuum on the parameters that matter.
Matrixing handles equipment families. For several equivalent units, or several sizes of the same design, validate one representative of each category and document why the others are equivalent: same material, same surface finish, same geometry, same cleaning recipe scaled by area.
Both have to be justified in the validation plan and approved by quality. A bracketing argument that skips a product because it is inconvenient, rather than because it is genuinely bracketed, is the kind of gap an inspector finds. The justification has to name the parameters being bracketed (solubility, ADE, dose, batch size) and show the chosen product sits at the demanding end of each, otherwise the “bracket” leaves a product uncovered.
Verification and Lifecycle Management
Routine cleaning verification. After each clean in production, confirm the clean was performed per the validated procedure. This is the operational check, cleaning time, agent concentration, temperature, rinse conductivity, and visual inspection, not full analytical residue testing on every batch. Analytical residue testing is the validation activity; verification is the day-to-day evidence that the validated state still holds. The records belong in the batch record review. The between-batch cleaning verification checklist turns this into a runnable, signed release gate, tied to the validated parameters, rather than a verbal “looked clean” judgment.
Periodic review and revalidation triggers. Cleaning validation is a living state, not a one-time certificate. Reassess or revalidate when any of these occur:
- A new product is introduced with a lower ADE, lower solubility, or larger surface footprint than the current worst case
- The cleaning procedure changes: agent, concentration, temperature, contact time, or rinse
- Equipment is modified in a way that affects product-contact surfaces, or a new piece is added to the train
- A product formulation changes in a way that affects how it cleans off
- Adverse trends appear in routine monitoring
Trending. Track swab and rinse results over time, not just pass or fail. A slow climb in residue, while still under the limit, can signal a degrading procedure: equipment wear and surface roughening, biofilm, scaling, or quiet procedure drift on the floor. Catching the trend before it crosses the limit turns a deviation into a maintenance action. The statistical tools for this, control charts and capability indices, are covered in statistics in quality.
For the hands-on execution side of all this, sampling logistics, run sequencing, deviation handling during the runs, and report assembly, see the companion article on cleaning validation execution.
Common Cleaning Validation Failures
No ADE. Using 10 ppm or visual-only limits with no toxicological basis. Since the EMA 2014 guideline this is not acceptable for shared-facility justification.
Swab recovery never established. Reporting a surface result without correcting for recovery overstates how clean the equipment is.
Wrong worst case. Picking the worst-case product on solubility when the limit should be driven by ADE, or the other way around. Solubility drives the cleaning strategy, ADE drives the acceptance limit. Confusing the two produces a limit that is either indefensible or impossibly tight.
Validation not maintained. New products added to the equipment without reassessing the matrix. The most damaging version is a high-potency compound introduced onto a shared train that was validated only against low-potency products.
Procedure not followed during validation. Validating with more rinses, longer contact, or hotter water than the SOP allows. The validation then proves something production never does.
No defined visual criteria. “No visible residue” is not objective unless the lighting, distance, angle, and a reference standard for “visible” are written down and trained.
Ignoring degradants. Cleaning agents can degrade an API into a different molecule, and detergents can leave their own degradation products. Validating for the parent API only can miss a residue the method does not even look for.
Analytical method not fit for the limit. A method whose limit of quantitation sits above the calculated acceptance limit cannot prove a pass. The result reads “less than LOQ” while the surface could still be over the limit. Confirm method sensitivity against the swab acceptance number before execution.
Hold times not validated. Running cleaning validation with a same-day clean while the SOP allows a 72-hour dirty-hold, or sampling immediately while production stores cleaned equipment for a week. The validated state has to bracket what production actually does.
No documented dedication decision. A highly potent or sensitizing product placed on shared equipment with cleaning validation as the only control, and no separate, approved quality risk management record showing why sharing was acceptable rather than dedicating.
Biologics residue tested with small-molecule methods only. Relying on TOC or a generic assay as the sole test for a protein-based process, with no HCP, DNA, or endotoxin-specific method addressing the residues that actually drive patient risk for that modality.
Single-use components treated as needing no assessment. Assuming a disposable bag or tubing set needs no evaluation because it is discarded after one use, when extractables and leachables into the product still require the same discipline as any other product-contact material.
What Inspectors Look For
Cleaning validation is a recurring inspection focus, in both FDA and EU inspections, and the requests are predictable. Inspectors ask for:
- The cleaning validation protocol and report, with the swab location maps
- The ADE or PDE monographs and the toxicological assessments behind them
- Swab recovery studies for each analyte and surface material
- Evidence that the worst-case product was correctly identified, with the reasoning
- The current product matrix for the equipment, and proof that newly introduced products were assessed against it
- Cleaning verification records for recent batches, to confirm the floor follows the validated procedure
- Confirmation that the procedure in use matches the validated procedure exactly
- A documented dedicated-versus-shared decision for any potent, sensitizing, or genotoxic product on the equipment, not just the cleaning validation package for products already deemed shareable
Finding 10 ppm limits with no ADE, swab recovery never established, or a potent new compound added to shared equipment without reassessment will draw an observation every time. For how those observations are written up and answered, see fda inspection readiness and the 483 response strategy.
Interview Questions and How to Answer Them
These come up in QA, validation, and manufacturing-science interviews, and the answers reveal quickly whether someone has actually run a cleaning validation or only read about it.
“Walk me through how you set a cleaning limit.” Start from the ADE of the worst-case previous product, calculate MACO using the next product’s batch size and maximum daily dose, divide by total shared product-contact surface area to get a µg/cm² limit, then carry that limit through the swab area, extraction volume, and recovery factor to get the acceptance concentration the lab reports against. Mention that you would cross-check against the legacy 0.1% dose and 10 ppm methods and take the most stringent. The interviewer is checking that you can chain ADE to MACO to surface limit to swab number without hand-waving.
“What is the difference between ADE and PDE?” Same concept, different name. ADE is the ISPE and US-practice term, PDE is the EMA term, and they produce the same value from the same toxicological derivation. A single derived value drives the limit in both regions.
“Which is the worst case, the product that is hardest to clean or the most toxic?” Both, on different axes. The hardest-to-clean product (poor solubility) defines the cleaning challenge you validate against. The most toxic product (lowest ADE) defines the acceptance limit. Best practice validates the harder soil but applies the tighter limit. Saying “it depends” without separating the two axes is the wrong answer.
“Why three runs?” It is the conventional minimum from the 1993 FDA guide and it demonstrates reproducibility, not a single lucky pass. They must be consecutive and at worst-case conditions, and a failure in the middle restarts the count. Add that modern lifecycle thinking treats three runs as the qualification stage, followed by continued verification.
“How do you handle swab recovery?” Spike known amounts onto coupons of each contact material, swab with the production technique, extract, and analyze; recovery should be at least 70% and consistent across the range. Correct every reported result by the recovery factor. Recovery below 50% is generally not acceptable. The follow-up they want: you run recovery on every surface material, not just stainless, because elastomers and PTFE recover differently.
“You add a new, more potent product to a validated shared line. What do you do?” Trigger a change control, derive its ADE, re-run the worst-case assessment, and very likely revalidate, because a lower-ADE product can drop the acceptance limit below what the existing procedure was shown to achieve. If the compound is highly sensitizing or cytotoxic, escalate the question of whether the line can be shared at all or needs dedication.
“What is the role of visual inspection?” A first-pass qualitative gate with defined lighting, distance, and angle, useful because it is immediate, but not the quantitative acceptance criterion. For potent compounds the calculated limit is below the visible threshold, so visual can pass while the surface is over the limit.
“When would you dedicate equipment instead of validating a shared cleaning process?” When the hazard classification alone rules it out, highly potent, sensitizing, or a non-threshold genotoxin, or when the ADE-driven acceptance limit sits below what the cleaning procedure and the analytical method can actually demonstrate. I would run the numbers, compare them to real demonstrated recovery-corrected performance and method LOQ, and document the decision and the rejected alternative in a quality risk management record approved by Quality, not just assume sharing is fine because a cleaning validation program exists.
“How does cleaning validation differ for a biologic compared to a small molecule?” The chain, limit, sample, recover, and trend, is the same, but the dominant residues change: host-cell protein, residual DNA, and endotoxin instead of a classic dose-response toxicant, so the assays change to ELISA, qPCR, and LAL rather than relying on TOC or HPLC alone. Endotoxin gets its own clean limit from the K-over-M threshold-pyrogenic-dose formula. And a lot of modern biologics manufacturing sidesteps the cross-contamination question entirely with single-use systems, where the validation question shifts to extractables and leachables rather than swab and rinse.
Cleaning Validation Program Readiness Checklist
Use this before declaring a shared-equipment cleaning validation program inspection-ready. Each item should have a real document behind it, not a verbal assurance.
| # | Check | Evidence expected |
|---|---|---|
| 1 | Every product on the shared train has a signed ADE/PDE (HBEL) monograph | Toxicology monograph, current, reviewed on schedule |
| 2 | The dedicated-versus-shared decision is documented for any potent, sensitizing, or genotoxic product | Quality risk management record with Quality approval |
| 3 | Worst-case product and worst-case equipment are selected with a scored, documented rationale | Worst-case matrix, not a narrative assertion |
| 4 | MACO, surface limit, and swab/rinse acceptance numbers are calculated and cross-checked against the 10 ppm and dose criteria | Completed MACO worksheet, most conservative limit carried forward |
| 5 | Swab and rinse recovery are established for every relevant surface material, not just stainless steel | Recovery study report, factor and %RSD per material |
| 6 | The analytical method’s limit of quantitation sits below the acceptance limit with margin | Method validation report referencing the specific limit |
| 7 | Sampling locations are numbered on equipment diagrams and justified as worst case | Sampling plan with diagrams or labeled photos |
| 8 | Dirty hold time and clean hold time are validated at their real maximums | Hold-time study data, not assumed defaults |
| 9 | The cleaning procedure in the SOP matches exactly what was validated | Side-by-side comparison, no undocumented gap |
| 10 | Three consecutive successful runs, or a justified alternative, are complete and documented | Executed protocol data, no failure buried mid-sequence |
| 11 | Cleaning agent residue is addressed on its own, not folded silently into the active’s result | Cleaning-agent limit and test result |
| 12 | The validation summary report states a clear conclusion and reconciles every deviation | Approved report, all deviations listed with impact |
| 13 | Routine cleaning verification and periodic review triggers are defined for life after validation | Verification records and a periodic review procedure |
| 14 | For biologics or single-use trains, biologics-specific assays (HCP, DNA, endotoxin) or extractables/leachables assessment are in place as applicable, not substituted with small-molecule methods alone | Assay validation records or E&L assessment |
Practical Tips
- Pre-calculate the swab acceptance concentration in the units the lab reports and put it in the protocol, so the analyst is comparing apples to apples and not back-converting at disposition.
- Photograph and number every swab location, and keep the map with the report. “We swabbed the worst spots” is not a map.
- Confirm the analytical method LOQ beats the limit before you generate any samples; discovering the method cannot reach the limit after execution wastes the whole study.
- Validate hold times at their maximum, not at convenience. The dirty-hold and clean-hold in the SOP must be the ones you proved.
- For dedicated equipment, do not skip cleaning validation entirely; shift the question to degradant and residue build-up over successive batches of the same product.
- Keep toxicology independent. The person who derives the ADE should not be under pressure from the schedule that needs the limit to be loose.
For the surrounding framework, the validation deliverables guide shows where the cleaning validation protocol and report sit among the other validation documents, and cross-contamination control in shared facilities covers the facility-design and dedication decisions that sit upstream of the cleaning limit.