This is a ready-to-use risk assessment for the two failure modes of an AI-based automated visual inspection (AVI) system: a false accept, a defective unit classified as good, and a false reject, a good unit classified as defective. The two carry different consequence classes, patient safety against yield and operational pressure, and scoring them into one averaged number hides the asymmetry that should drive every acceptance criterion and every mitigation on the system. This is narrower than a general AI/ML system risk assessment; use it alongside one to size the broader data-quality, drift, and autonomy factors of the system as a whole. Replace every <<FILL: ...>> placeholder with your own specifics, and route it through your normal risk-management and validation review. A worked filled specimen follows. Confirm each cited reference against the current source before you rely on it.
Document control header
| Field | Entry |
|---|---|
| Document title | AI-Based Automated Visual Inspection Misclassification Risk Assessment |
| Document number | <<FILL: RA-ID>> |
| Version | <<FILL: version>> |
| AVI machine, product, container in scope | <<FILL>> |
| Author | <<FILL: name, role>> |
| QA approver | <<FILL: name, role>> |
1. Methodology
Use a risk-based approach consistent with ICH Q9, scoring severity, occurrence, and detectability separately for each defect class and each failure direction, false accept and false reject, then combining into a risk class. Because a false accept and a false reject differ in kind, not just in degree, do not average them into one number; carry both through the assessment and the acceptance decision.
2. Scope
Applies to <<FILL: AVI machine ID, product, container>> at its qualified detection capability and imaging conditions. Assumes the imaging chain and the classifier have completed their own qualification (see the classifier validation protocol) and this assessment sizes the residual risk of the qualified system in routine operation, including the drift that can occur between periodic re-challenges.
3. Scoring scales
3.1 Severity, scored by failure direction
| Score | False accept consequence | False reject consequence |
|---|---|---|
| 5, catastrophic | An undetected critical particulate or defect reaches a patient lot with a plausible harm pathway | Not applicable to this direction: a false reject cannot itself harm a patient |
| 4, major | An undetected major defect (a non-critical particulate, a container-integrity concern) ships | A sustained false-reject rate high enough to cause a supply interruption |
| 3, moderate | An undetected minor cosmetic defect ships | An elevated false-reject rate causing yield loss and rework pressure |
| 2, minor | Folds into moderate for practical scoring | An occasional false reject, absorbed as normal operating yield |
| 1, negligible | Not applicable | Negligible operational effect |
3.2 Occurrence, the likelihood the failure happens at a materially higher-than-qualified rate
| Score | Descriptor | Example driver |
|---|---|---|
| 5 | Frequent | No monitoring, unstable imaging conditions, frequent uncontrolled retrains |
| 3 | Occasional | Known drift drivers exist (lamp aging, seasonal container tint) with only periodic re-challenge |
| 1 | Rare | Imaging conditions well controlled, monitoring continuous, change control mature |
3.3 Detectability, how likely the organization is to catch the failure before it causes harm or loss
| Score | Descriptor | Example |
|---|---|---|
| 1 | Very likely to catch | Independent AQL sampling in place, continuous per-class reject-mix trending, a genuinely exercised human pre-sort step |
| 3 | Possible | Aggregate-only reject-rate trending, periodic re-challenge only |
| 5 | Unlikely to catch | No independent backstop, no monitoring, a fully automated decision with no human check anywhere in the chain |
4. Risk assessment table
| Defect class | Failure direction | Severity | Occurrence | Detectability | Risk score | Risk class |
|---|---|---|---|---|---|---|
<<FILL: critical particulate>> | False accept | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> |
<<FILL: critical particulate>> | False reject | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> |
<<FILL: cosmetic glass defect>> | False accept | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> |
<<FILL: cosmetic glass defect>> | False reject | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> |
5. AI-specific risk factors to weigh into occurrence and detectability
- Drift exposure. How many known drift drivers (lighting, lamp aging, container or product changes) act on this line, and how tightly are they controlled.
- Training-data representativeness. Whether the class’s reference and training units span its real severity and size range, including the gray zone, or only the easy cases.
- Catalog completeness. Whether a defect type outside the current catalog could plausibly occur and go entirely unclassified by any control.
- Automation bias. For a human pre-sort step, whether the override or disagreement rate is monitored, since a decayed pre-sort silently removes a detectability control this assessment may be counting on.
- Change-control maturity. Whether model, recipe, and imaging changes are governed by a predetermined change control plan or handled ad hoc.
6. Mitigations and residual risk
For each unacceptable risk class, define the mitigation, the resulting occurrence or detectability change, and the residual risk.
| Risk (from section 4) | Mitigation | Effect | Residual risk class |
|---|---|---|---|
<<FILL>> | <<FILL: e.g. add continuous per-class reject-mix trending>> | Detectability improves from 3 to 1 | <<FILL>> |
<<FILL>> | <<FILL: e.g. tighten imaging-condition monitoring cadence>> | Occurrence improves from 3 to 1 | <<FILL>> |
State explicitly which mitigations rely on the independent AQL manual sampling step and the qualified manual fallback, since both act as risk-reducing controls that exist outside the AI system itself and should be credited as such in this assessment, not assumed silently.
7. Approval
| Role | Name | Signature | Date |
|---|---|---|---|
| Author | <<FILL>> | ||
| Inspection SME | <<FILL>> | ||
| QA | <<FILL>> |
8. References
ICH Q9, Quality Risk Management. USP General Chapter <1790>, Visual Inspection of Injections. USP General Chapter <790>, Visible Particulates in Injections. EU GMP Annex 1 (2022).
Confirm the current version of each reference before issue.
Revision history
| Version | Date | Author | Summary of change |
|---|---|---|---|
<<FILL: 1.0>> | <<FILL: date>> | <<FILL: author>> | Initial issue. |
Filled specimen
Illustrative scoring for two defect classes on a vial-line AVI system with a fully automated accept/reject decision and no human pre-sort step, continuous per-class reject-mix trending, and quarterly imaging-condition checks.
| Defect class | Failure direction | Severity | Occurrence | Detectability | Risk score | Risk class |
|---|---|---|---|---|---|---|
| Critical particulate (100-150 um) | False accept | 5 | 3 (lamp aging is a known driver, periodic re-challenge only at release) | 1 (AQL sampling and continuous per-class trending in place) | 15 | High, mitigate |
| Critical particulate | False reject | 2 | 3 | 3 | 18 | Moderate, monitor |
| Cosmetic glass defect | False accept | 3 | 3 | 3 (aggregate reject trending only for this class, no dedicated per-class monitor) | 27 | High, mitigate |
| Cosmetic glass defect | False reject | 3 | 2 | 1 | 6 | Low, accept |
Mitigation applied: continuous per-class reject-mix trending was extended to the cosmetic-glass class (previously trended only in the aggregate reject rate), moving its detectability from 3 to 1 and its risk score from 27 to 9, moderate. The critical-particulate false-accept risk was mitigated by tightening the re-challenge cadence from quarterly to monthly for that class specifically, moving occurrence from 3 to 1 and the risk score from 15 to 5, low. Residual risk for both false-accept rows is now judged acceptable given the independent AQL sampling backstop, which the assessment credits explicitly rather than assuming.
Common inspection findings this risk assessment prevents
- One blended risk score across false accept and false reject, hiding the patient-safety-critical direction inside an average with the yield-only direction.
- A risk assessment that never mentions the AQL sampling step or the human pre-sort step as a detectability control, so their contribution is assumed rather than credited and reviewed.
- Occurrence scored the same for every defect class regardless of known, documented drift drivers specific to that class.
- A “mitigated” risk with no re-scored detectability or occurrence to show the mitigation actually changed the number.
How to adapt this risk assessment
- Set your own severity, occurrence, and detectability scales and thresholds in section 3, calibrated to your product’s actual harm pathway and your program’s actual monitoring maturity.
- Populate section 4 for every defect class in your catalog, in both failure directions, not just the critical ones; a class with only a false-reject risk still needs the row so the omission is a decision, not an accident.
- Tie each mitigation in section 6 to a specific, re-scoreable change, such as a new monitoring signal or a tightened re-challenge cadence, rather than a general statement of intent.
- Reference the classifier validation protocol and the monitoring log this assessment depends on for its detectability assumptions.
- Confirm each reference in section 8 against its current published version before issue.