Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
Risk Assessment Plug-and-play starting point AI & Automation

Risk Assessment: Generative AI Assistant in Deviation, CAPA, and Investigation Workflows

A plug-and-play FMEA-style risk assessment for a generative-AI drafting assistant in quality operations: the failure modes specific to generative models (confabulation, miscounting, generic CAPAs, data leakage), scoring scales, mitigations, and residual risk.

Document type: Risk Assessment

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use risk assessment for a generative-AI assistant used to draft and summarize inside deviation, CAPA, and investigation workflows. It uses a failure-mode-and-effects (FMEA) structure and concentrates on the failure modes a deterministic-software assessment never had to consider. Replace every <<FILL: ...>> placeholder with your own specifics and adjust the scoring to your own scales. This content is educational and general; adapt it and verify it before use.

Assessment header

FieldEntry
System / assistant<<FILL: assistant name / ID>>
Intended useDrafting and summarizing in deviation, CAPA, and investigation workflows, with mandatory human review; no autonomous quality decision
Assessment ID<<FILL: RA-ID>>
Assessors<<FILL: QA, System Owner, Data Science, IT>>
Date<<FILL: date>>
Risk methodFMEA (Severity x Occurrence x Detection), per <<FILL: SOP-ID for quality risk management>> and ICH Q9(R1)

1. Methodology

Each failure mode is scored for Severity (impact on product quality, patient safety, or record integrity if it reaches a record undetected), Occurrence (how often the failure mode is expected to arise), and Detection (how likely the human-review control is to catch it before the record is finalized, where a high score means poor detectability). The Risk Priority Number (RPN) is Severity x Occurrence x Detection. The human-review step is treated as the primary detection control throughout, because the whole approach depends on it. Scoring the failure modes as if review did not exist would overstate residual risk; scoring them assuming perfect review would understate it, so Detection is scored on how well a realistic, trained reviewer catches each mode.

2. Scoring scales

Severity (1-5)

ScoreMeaning
5An undetected error could directly affect product disposition, patient safety, or a regulatory submission
4Could produce an inaccurate GxP record used in a quality decision
3Could weaken an investigation or CAPA without an immediate product impact
2Minor record-quality impact, correctable at next review
1Negligible

Occurrence (1-5)

ScoreMeaning
5Expected in routine use without a specific control
3Occasional
1Rare

Detection (1-5, higher = harder to detect)

ScoreMeaning
5A trained reviewer would usually miss it (for example a fluent, plausible fabrication)
3A trained reviewer catches it with defined checks
1Obvious; almost always caught

3. The assessment

#Failure modeEffectCauseSODRPNMitigationResidual S/O/DResidual RPN
1Confabulation: a fluent, confident, factually wrong statement in a draftInaccurate GxP record if adopted; investigation misdirectedModel draws on training data instead of supplied facts44580Grounding to supplied facts; fact-by-fact verification in review (WI step 2); disclosure flag4 / 2 / 216
2Miscounted or invented number in a summaryInaccurate trend record; false or missed signalLanguage models are unreliable at counting/arithmetic over large inputs44464Deterministic computation of counts; verify every number against source (WI step 3)4 / 2 / 216
3Premature cause in a deviation descriptionInvestigation anchored on a wrong cause before it startsModel speculates when not constrained to observations33327Prompt constrains to observations; reviewer checks for prejudged cause (WI step 4)3 / 1 / 26
4Generic reflex CAPA (blanket retraining, “update the SOP”) not tracing to the causeIneffective CAPA; recurrence; repeat becomes its own findingModel defaults to plausible generic actions44348Reviewer rejects any action not traced to the confirmed root cause (WI step 4)4 / 2 / 216
5Model’s suggested cause accepted as the answerRoot cause effectively decided by the toolAutomation bias amplified by fluent output53460Human evidence-weighing mandatory; cause attributed to named investigators; RCA never delegated5 / 1 / 210
6Fabricated “similar past deviation” citationInvestigation relies on a comparison that does not existModel invents a plausible reference33436Every cited historical record confirmed to exist and be relevant before use3 / 1 / 26
7Confidential GxP content entered into a public AI toolData leaves control; confidentiality and integrity eventStaff use an unsanctioned tool for convenience43448Sanctioned contained tool only; trained prohibition; monitoring/DLP where available4 / 1 / 312
8Silent truncation of a long inputSummary misses records; a real signal is buriedInput exceeds the tool’s context window33436Reviewer confirms coverage; chunking with reconciliation; completeness check (WI step 5)3 / 2 / 212
9Vendor model change alters behavior in productionValidated state lapses silently; output quality shiftsVendor updates the base model with no version bump on your side43448Version pinning where available; vendor change treated as change control; re-confirm guardrails4 / 2 / 216
10Rubber-stamp reviewThe control the approach depends on is not actually performedReviewer approves polished drafts without engaging53460Meaningful-review WI with per-step acceptance; edit capture; QA monitors edit-rate as a signal5 / 2 / 220
11Undisclosed AI use discovered in inspectionFinding for hidden, unassessed AI in a quality processNo disclosure flag or procedure42324Mandatory AI-assistance flag; procedure defines permitted use; transparency in the record4 / 1 / 28
12No record of generation conditionsBasis of a draft cannot be reconstructedModel version and prompt not captured23318Capture model version and prompt/template version on the disclosure log where feasible2 / 1 / 24

4. Risk acceptance

Set an action threshold appropriate to your scales (for example, any residual RPN above <<FILL: threshold>>, or any residual Severity of 5 with Occurrence above 1, requires a further control or management sign-off). The two residual risks that stay highest in this assessment are rubber-stamp review (#10) and model-decided root cause (#5), both because their severity is inherent and the mitigation is human discipline rather than a code control. Manage them with the meaningful-review work instruction, reviewer training on these exact failure modes, and QA monitoring of the draft edit-rate as a leading indicator that review has gone shallow.

5. Residual risk statement

With grounding, mandatory fact-and-number verification, a trained meaningful-review control, a sanctioned contained tool, disclosure flagging, and change control over vendor model updates in place, the residual risk of the assistant in its stated intended use is <<FILL: acceptable / acceptable with monitoring>>. The approach is defensible because the assistant only drafts and summarizes, a named human owns every decision and record, and the highest residual risks are actively monitored rather than assumed away.

6. Approval

RoleNameSignatureDate
Assessment author<<FILL>>
Quality Assurance<<FILL>>
System Owner<<FILL>>

Common inspection findings this assessment addresses

  • No documented risk assessment for an AI tool in use in quality workflows.
  • A risk assessment that treats the assistant like deterministic software and never names confabulation, miscounting, or generic-CAPA failure modes.
  • Mitigations listed but not tied to a specific, verifiable control (a work-instruction step, a disclosure field, a change-control trigger).
  • No leading indicator (such as edit-rate monitoring) for the review control the whole approach depends on.

How to adapt this assessment

  1. Replace the scoring scales with your own quality-risk-management scales and re-score to them.
  2. Add or remove failure modes for the specific workflows and tool you deploy.
  3. Tie each mitigation to your real work instruction, disclosure log, and change-control procedure.
  4. Set your action threshold and route residual risks above it for management acceptance.
  5. Re-run the assessment on a vendor model change or an intended-use change.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.