Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
Protocol Plug-and-play starting point AI & Automation

Protocol: AI/ML Classifier Validation for Automated Visual Inspection

A plug-and-play validation protocol for the trained AI or machine learning classifier inside an automated visual inspection machine: imaging-chain preconditions, a locked probability-of-detection challenge set, per-class sensitivity, specificity, and false-reject-rate acceptance criteria, and a filled specimen, distinct from human-inspector qualification and from a rule-based-versus-manual comparison.

Document type: Protocol

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use validation protocol for the trained AI or machine learning classifier inside an automated visual inspection (AVI) machine for injectable products. It is a different exercise from qualifying a human inspector, which judges a different population (people, not a frozen artifact) against effectiveness and efficiency criteria, and from comparing a rule-based automated machine to manual inspection by the Knapp method, which compares a deterministic recipe to a human benchmark. Here the subject is a trained model: its capability is proven on a locked, held-out image test set against per-class detection, specificity, and false-reject acceptance criteria set before training, and its qualified state holds only within the imaging conditions under which that capability was measured. Replace every <<FILL: ...>> placeholder with your own specifics, set your document numbers and dates, and route it through your normal validation review and approval. A worked filled specimen follows the template. Verify each cited regulation against the current source before you rely on it.

Approval page

FieldEntry
Protocol titleAI/ML Classifier Validation for Automated Visual Inspection
Protocol number<<FILL: protocol ID>>
AVI machine, model, and version<<FILL: machine ID, model name, model version>>
Product and container in scope<<FILL: product, container format>>
Author<<FILL: name, role>>
QA / Validation approver<<FILL: name, role>>
Data Science / ML Engineering approver<<FILL: name, role>>

1. Objective

Demonstrate that the trained AI classifier behind the AVI machine detects each defect type in its catalog at a stated, defined probability, measured on a locked image test set the model never saw during training or tuning, under imaging conditions that are separately qualified, and that its false-reject rate on good units stays within an operationally acceptable limit.

2. Scope

This protocol qualifies the classifier’s detection performance for the named product-and-container combination and the named defect catalog. It assumes:

  • The imaging chain (optics, lighting, container handling, camera and trigger stability) has completed its own qualification per <<FILL: reference to imaging IQ/OQ>>, and this protocol’s capability claim is valid only within those qualified conditions.
  • The equipment and computerized-system controls (audit trail, access control, data integrity of captured images and decision records) are covered by <<FILL: reference to EQ/CSV documents>>.
  • This protocol does not qualify a human pre-sort or review step, which is qualified separately per <<FILL: reference, e.g. inspector qualification procedure>>, and it does not replace the manual AQL acceptance-sampling step required under USP <790>.

3. System description

<<FILL: describe the model architecture family, the training framework, the deployment platform, and how the model's output interfaces with the imaging and rejector hardware>>

4. Prerequisites

  • Imaging chain qualification is complete and current, and its qualified conditions (optics, lighting, container and product presentation, camera and trigger parameters) are recorded in <<FILL: reference>>.
  • The defect catalog is approved, naming each defect type, its size or severity range, and its criticality.
  • The training, tuning, and test image datasets are built, characterized, labeled by qualified inspectors with measured inter-labeler agreement, and split into three, with the test set locked and never used in training or tuning. Every dataset version is hashed and recorded.
  • Detection-capability acceptance criteria (section 6) were written and approved before the model was trained or tested against the locked set.

5. Roles

RoleResponsibility
Validation leadExecutes the protocol, records results, dispositions deviations.
Data Science / ML EngineeringProvides the trained model at a fixed version, the dataset lineage and version hashes, and executes the locked-set run under witness.
Inspection SMEConfirms the defect catalog, the reference-unit characterization, and the labeling used to build the challenge set.
QA / ValidationApproves acceptance criteria before training, witnesses the locked-set run, and approves the final qualification decision.

6. Acceptance criteria

Set per defect class, before training. State the metric, the target, the justification tied to defect criticality, and the population and conditions the target must be met on.

Defect classCriticalityDetection probability (sensitivity) targetFalse-reject ceiling on good unitsBasis
<<FILL: e.g. critical particulate, size band>>Critical>= <<FILL: e.g. 0.95>><= <<FILL: e.g. 3%>><<FILL: patient-safety rationale>>
<<FILL: e.g. cosmetic glass defect>>Major>= <<FILL>><= <<FILL>><<FILL>>
<<FILL: e.g. fill-level fault>>Major>= <<FILL>><= <<FILL>><<FILL>>
Good units (specificity)Not applicableSpecificity >= <<FILL: e.g. 0.97>>Not applicableYield and operability

For a class with fewer than <<FILL: e.g. 30>> characterized units in the locked test set, report a 95 percent confidence interval (for example a Wilson score interval) alongside the point estimate, and do not accept a class on a point estimate alone. A wide interval on a rare class is itself an acceptance-relevant fact, not a footnote.

7. Procedure and test cases

7.1 Confirm imaging preconditions (Test case C1)

ItemDetail
StepConfirm imaging chain qualification is current and the recipe version matches the version this protocol tests
ExpectedQualified and current; recipe version recorded
Actual / Pass-Fail<<FILL>>

7.2 Confirm dataset integrity (Test case C2)

ItemDetail
StepConfirm training, tuning, and locked test datasets are version-hashed, the test set was never used in training or tuning, and labeling agreement is on file
ExpectedVersion hashes recorded; test-set isolation confirmed; agreement documented
Actual / Pass-Fail<<FILL>>

7.3 Run the locked test set (Test case C3)

  1. Run the frozen model, at its recorded version, against the entire locked test set once, under witness, with no further tuning permitted after the run begins.
  2. Record, per unit, the ground-truth label, the model’s decision and class, and the confidence score.
  3. Build the confusion matrix per defect class.
ItemDetail
Test-set composition<<FILL: units per class, natural vs characterized/seeded, good-unit count>>
ExpectedA complete decision record for every unit; no re-runs after the witnessed start
Actual / Pass-Fail<<FILL>>

7.4 Compute and compare per-class metrics (Test case C4)

For each defect class, compute:

  • Sensitivity (probability of detection) = true positives / (true positives + false negatives).
  • False-reject rate on good units = false positives / (false positives + true negatives); specificity = 1 minus this rate.
  • A 95 percent confidence interval where the class has fewer characterized units than the threshold set in section 6.

Compare each result to its section 6 target.

Defect classTPFNFP (on good units)TNSensitivitySpecificity / false-reject rateTarget met?
<<FILL>><<FILL>><<FILL>><<FILL>><<FILL>><<FILL>><<FILL>><<FILL>>

7.5 Confirm the imaging-condition boundary (Test case C5)

State explicitly, in the summary, that the reported capability holds only within the qualified imaging conditions recorded in section 4, and that any change to those conditions requires re-qualification before the capability claim can be relied on again.

8. Deviation handling

Any departure from this protocol during execution, including a re-run of the locked test set for any reason, is recorded as a protocol deviation, assessed for impact on the qualification decision, and dispositioned before qualified status is granted. Reusing a locked test set after a failed run with no documented, QA-approved rationale invalidates the claim that the set was ever locked.

9. Summary and conclusion

FieldEntry
Per-class results vs targets<<FILL: pass/fail per class>>
Classes carrying a confidence interval<<FILL>>
Imaging conditions the result is scoped to<<FILL>>
Overall resultQualified / Not qualified
Scope<<FILL: machine, model version, product, container, recipe version>>
Approved by (name, date)<<FILL>>

10. Attachments

  • Defect catalog and criticality assignment.
  • Dataset lineage record (training, tuning, and locked test sets, versions, labeling procedure and measured agreement).
  • Confusion matrices and confidence intervals per class.
  • Imaging chain qualification reference.

11. References

USP General Chapter <1790>, Visual Inspection of Injections (probability of detection, capability concepts). USP General Chapter <790>, Visible Particulates in Injections. EU GMP Annex 1 (2022), on validation of automated inspection equipment. ICH Q9, Quality Risk Management, for sizing acceptance criteria to defect criticality. GAMP 5 (2nd edition) and your computer software assurance procedure, for the software-validation lifecycle this protocol sits inside.

Confirm the current version and clause numbers of each reference before issue.

Revision history

VersionDateAuthorSummary of change
<<FILL: 1.0>><<FILL: date>><<FILL: author>>Initial issue.

Approvals

RoleNameSignatureDate
Author<<FILL>>
Data Science / ML Engineering<<FILL>>
QA / Validation<<FILL>>

Filled specimen

An illustrative qualification run for a vial-line AVI classifier, model version clf-v3.2, recipe version 4. Numbers are examples.

Defect classTPFNFPTNSensitivitySpecificity / FRRTarget met?
Critical particulate (100-150 um)582187820.967 (95% CI 0.886-0.991), target >= 0.95FRR 2.25%, target <= 3%Yes
Cosmetic glass defect355227780.875 (95% CI 0.739-0.945), target >= 0.90FRR 2.75%, target <= 3%No, sensitivity below target
Fill-level fault24197910.96 (95% CI 0.805-0.993), target >= 0.90FRR 1.13%, target <= 3%Yes, small n, interval carried

Decision: not qualified as submitted. The critical particulate and fill-level classes met their targets, the fill-level result carrying its confidence interval because the class had only 25 characterized units. The cosmetic glass class fell short of its 0.90 sensitivity target at 0.875, with a confidence interval (0.739 to 0.945) that does not clearly clear the target either. Deviation raised; Data Science added characterized cosmetic-defect units concentrated on the gray zone and retrained; the class was re-tested on an expanded locked set at the next protocol execution before qualified status was granted for the full catalog. The imaging conditions (recipe version 4, qualified lighting and optics per <<FILL>>) are the boundary the qualified capability is scoped to; a lamp, lens, or handling change invalidates this result until re-qualification.

Common inspection findings this protocol prevents

  • A single blended accuracy number reported for the whole catalog, hiding a weak class like the cosmetic-glass example above.
  • A capability number with no stated imaging conditions, so it cannot be tied to whether the running line still matches the tested state.
  • A rare defect class accepted on a raw point estimate with no confidence interval, overstating what a handful of units actually proves.
  • Test images that were also used in training or tuning, inflating the reported capability.
  • A classifier compared only to itself, with no record that acceptance criteria were written before the model was trained and tested.

How to adapt this protocol

  1. Set the defect catalog, the criticality of each class, and the acceptance targets in section 6 from your own risk assessment, before training begins.
  2. Size the locked test set per class so the confidence interval is acceptably narrow for a critical class; where natural defects are too rare, use characterized or seeded units and document how each was created.
  3. Point sections 2 and 4 at your real imaging-chain qualification, equipment qualification, and computerized-system validation documents.
  4. Reference your predetermined change control plan for how a failed class (as in the specimen) is remediated and re-tested.
  5. Confirm each reference in section 11 against its current published version before issue.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.