Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
SOP Plug-and-play starting point AI & Automation

SOP: Human-in-the-Loop Review and Rejection Sampling for AI-Assisted Case Processing

A plug-and-play SOP for the human review control over AI-assisted pharmacovigilance case processing: what the reviewer confirms for a human-confirmed model, the rejection-sampling program that catches what a screening model silently drops, automation-bias monitoring, and escalation, with a filled specimen and the regulations it satisfies.

Document type: SOP

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use SOP for the human control that stands over an AI-assisted step in pharmacovigilance case processing: intake triage, MedDRA coding suggestion, literature screening, duplicate detection, narrative drafting, or signal-detection augmentation. It defines what a reviewer confirms for a model a human confirms on every output, and, separately, the rejection-sampling program required for a model that gates what a human sees at all. It is meant to sit alongside, not replace, SOP: ICSR intake, triage, and processing and SOP: signal detection and management, adding the AI-specific review control those procedures do not cover. Replace every <<FILL: ...>> placeholder, route it through document control, and confirm every cited regulation against the current source before you rely on it. This content is educational reference, not legal or regulatory advice.

Document control header

FieldEntry
Document titleHuman-in-the-Loop Review and Rejection Sampling for AI-Assisted Case Processing
Document number<<FILL: SOP-ID, e.g. SOP-PV-AI-006>>
Version<<FILL: version, e.g. 1.0>>
Effective date<<FILL: effective date>>
Supersedes<<FILL: prior version or "New">>
Document owner<<FILL: role, e.g. Head of Pharmacovigilance>>
Applies to<<FILL: products / regions / AI-assisted processes in scope>>

1. Purpose

This procedure defines how <<FILL: COMPANY NAME>> maintains a meaningful human control over every AI-assisted step of case processing, so that a model output never reaches a case record, a submission, or a signal disposition without a qualified person exercising documented judgment over it, and so that a model that filters or screens what a human sees is checked for what it silently drops.

2. Scope

This procedure applies to every AI or machine learning model in production use anywhere in the pharmacovigilance workflow of <<FILL: COMPANY NAME>>, including models embedded in a commercial safety database platform. It covers the human-review step for a model a human confirms on every output (risk pattern A), the rejection-sampling program for a model that gates what a human sees (risk pattern B), and the escalation route when either control finds a problem. It does not cover model validation, governed by <<FILL: PRT-ID for AI/ML validation>>, or the underlying case-processing and signal-management procedures, governed by <<FILL: SOP-IDs>>, which this procedure supplements rather than replaces.

3. Responsibilities

RoleResponsibility
Reviewer (case processor, coder, safety physician, as applicable)Reviews every in-scope model output, confirms or overrides with a recorded reason, and escalates a pattern of concern.
PV / Safety System OwnerOwns the AI register entry for each in-scope model, its assigned risk pattern, and the rejection-sampling program for any Pattern B model.
PV Quality / QAAudits review records for evidence of meaningful review, not rubber-stamping; owns the automation-bias monitoring metric.
Rejection-sampling reviewerA qualified person independent of the routine review, who periodically re-reviews a sample of a Pattern B model’s rejections against ground truth.
QPPV or delegateInformed of any rejection-sampling finding that indicates a missed reportable case or a missed signal-relevant article.

4. Definitions

  • Human-confirmed assistance (risk pattern A): a model whose every output is reviewed by a qualified person before it has any effect on a case record, submission, or signal disposition.
  • Model-gated screening (risk pattern B): a model that decides what a human sees, for example by discarding messages classified “not relevant” or filtering literature before it reaches a screener. The dangerous failure is the false negative that never reaches a reviewer.
  • Rejection sampling: a periodic, independent human re-review of a representative sample of what a Pattern B model did not surface, used to estimate the false-negative rate the routine workflow cannot see.
  • Automation bias: the tendency of a reviewer to accept a usually-correct model’s suggestion without genuinely checking it, which hollows out the human control even when the review step is nominally performed.
  • Meaningful review: a review in which the reviewer sees the model output and enough context to judge it, applies independent judgment, and records a decision with a reason, as distinct from a review that only records that a screen was clicked.

5. Procedure

5.1 Determine the review design from the risk pattern

  1. For each AI-assisted process step, confirm its risk pattern from the current risk assessment (<<FILL: RA-ID>>). Do not assume; confirm the actual workflow matches the assumed pattern.
  2. For a Pattern A model, design the review per section 5.2. For a Pattern B model, design both the routine review of surfaced items per section 5.2 and the rejection-sampling program per section 5.3. A Pattern C model is out of scope for this SOP; do not deploy a model as Pattern C without the separate governance-board approval and QPPV sign-off that risk pattern requires.

5.2 Perform the routine human review

  1. Present the reviewer with the model output, its confidence or rationale where available, and the source material (verbatim, literature passage, cited cases) the output is based on.
  2. The reviewer confirms or overrides the output. A confirmation and an override are equally easy actions in the workflow; neither is the path of least resistance.
  3. Record a reason for every disposition, including a confirmation, not only an override. A one-line reason is sufficient; a blank reason field is not acceptable.
  4. Route a confirmed output that indicates a real problem (a valid case, a reportable article, a genuine signal pattern) into the governing procedure (<<FILL: SOP-ID>>) exactly as if a human had found it unaided.
  5. For a serious or high-consequence item (a possible fatal or life-threatening case, a signal-relevant cluster), require a second qualified reviewer or an escalation to a safety physician before the disposition is final.

5.3 Run the rejection-sampling program (Pattern B models only)

  1. Draw a representative sample of the model’s rejections (items it did not surface to the routine reviewer) on the cadence in the register entry (<<FILL: e.g. weekly for high-volume intake, monthly for literature screening>>), sized to detect a false-negative rate meaningful to the process (<<FILL: target detectable rate and confidence, e.g. detect a 5% miss rate at 90% confidence>>).
  2. Assign the sample to a rejection-sampling reviewer independent of the routine review queue, so a systemic gap is not checked by the same judgment that may have missed it.
  3. The rejection-sampling reviewer assesses each sampled item against ground truth, blind to the model’s original disposition where feasible.
  4. Compute the false-negative rate for the sample and compare it to the threshold in the register entry.
  5. Where a true miss is found, route it into the governing procedure immediately, independent of the sampling cadence; do not wait for the cycle to close to act on an individual missed case.

5.4 Monitor for automation bias

  1. Track the rate at which reviewers accept model outputs unmodified, by reviewer and in aggregate, on a rolling basis.
  2. Treat a rate near 100 percent sustained over time as a signal to investigate, not as evidence the model is excellent; a near-total acceptance rate as often means the review has become a rubber stamp as it means the model is right.
  3. Where investigation finds rubber-stamping, retrain the affected reviewers on the model’s known weaknesses and, where warranted, add a friction step (a required short justification for accepting a high-confidence suggestion on a serious item).

5.5 Escalate and record

  1. Escalate any rejection-sampling miss, any confirmed false-negative pattern, or any sustained automation-bias signal to the PV / Safety System Owner and QA the same working day.
  2. Open a deviation or investigation per <<FILL: SOP-ID for deviations>> where the finding could affect a reporting-clock case or a signal disposition.
  3. Record the finding and its resolution in the AI monitoring record (<<FILL: reference>>), separate from the individual case record it may also affect.

6. Acceptance criteria

  • Every AI-assisted process step has a review design that matches its confirmed risk pattern, not an assumed one.
  • Every reviewed output has a recorded disposition and reason, for both confirmations and overrides.
  • Every Pattern B model has a running, sized rejection-sampling program, and every sampled true miss is routed into the governing procedure and recorded.
  • Automation-bias monitoring is active for every in-scope model, with a defined response to a sustained high acceptance rate.
  • Every escalation reaches the PV / Safety System Owner and QA the same working day, and safety-relevant findings reach the QPPV or delegate.

7. References

EU GVP Module VI (case management) and Module IX (signal management). FDA draft guidance, “Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products” (issued 6 January 2025); draft, confirm current status. FDA and EMA, “Guiding Principles of Good AI Practice in Drug Development” (published jointly 14 January 2026), foundational, non-binding, human-centric. EMA, “Reflection paper on the use of artificial intelligence in the lifecycle of medicines” (adopted September 2024), non-binding. 21 CFR Part 11 and EU GMP Annex 11, for the review record as an electronic GxP record. ICH Q9(R1), Quality Risk Management, for sizing the rejection-sampling program to risk.

Confirm the current version and status of each reference before issue.

8. Record generated: AI-assisted review record

FieldEntry
Model / process<<FILL>>
Risk pattern<<FILL: A / B>>
Item reference (case ID, article ID, cluster ID)<<FILL>>
Model output<<FILL>>
Reviewer dispositionConfirm / Override / Needs investigation
Reason recorded<<FILL>>
Second-reviewer or escalation triggered<<FILL: Yes/No, and to whom>>
Rejection-sampling result (if applicable, per cycle)<<FILL: sample size, false-negative count and rate, threshold, pass/fail>>
Reviewer (name, date)<<FILL>>
QA review (name, date)<<FILL>>

9. Revision history

VersionDateAuthorSummary of change
<<FILL: 1.0>><<FILL: date>><<FILL: author>>Initial issue.

10. Approvals

RoleNameSignatureDate
Author<<FILL>>
Reviewer (QA)<<FILL>>
Approver (QPPV / PV Head)<<FILL>>

Filled specimen

A completed rejection-sampling cycle and one routine review entry for an example intake-triage model (Pattern B). Illustrative only.

FieldEntry
Model / processIntake triage assistant v1.4, inbound message classification
Risk patternB, model-gated screening
Rejection-sampling result (July 2026 cycle)Sample of 300 “not relevant” messages drawn from 6,140 total rejections; independent reviewer J. Alavi found 4 true misses (possible AE messages incorrectly discarded); false-negative rate 1.3%, within the 5% threshold
Escalation triggeredYes. All 4 missed cases routed to case intake same day; day zero re-established from the original message date, not the discovery date; 3 of 4 were within their reporting window despite the delay, 1 required a late-report assessment under the deviation procedure
Routine review entry (example)Message MSG-2026-08812, model flagged “possible AE,” reviewer S. Marchetti confirmed, reason: “verbatim describes new-onset rash after third dose, consistent with a reportable event,” routed to case intake
Reviewer / QAS. Marchetti (routine review), J. Alavi (sampling reviewer), 04 August 2026; QA review R. Odusanya, 05 August 2026

The one late report found in this cycle is exactly why the rejection-sampling program exists: the routine review only ever sees what the model surfaces, so the 4 missed cases would never have appeared to anyone reviewing the queue alone. The independent sample found them, the day-zero determination was made from the original message date rather than the discovery date, and the case that fell outside its window went through a late-report assessment rather than being quietly backdated.

Common inspection findings this SOP prevents

  • A screening model running for months with no rejection-sampling program, so the false-negative rate is unknown and unmeasured.
  • A review step recorded for every model output, but with a near-100-percent acceptance rate that no one investigated as a possible rubber stamp.
  • A rejection-sampling miss found and corrected quietly with no deviation opened, when the miss could have affected a reporting-clock case.
  • A reviewer disposition recorded with no reason, so a later investigator cannot tell whether the output was actually checked.
  • A risk pattern assumed from the design document rather than confirmed against the workflow actually running.

How to adapt this SOP

  1. Set your document number, owner, and effective date in the header.
  2. List every actual AI-assisted process step in scope, with its confirmed risk pattern, in section 2 and the register you point to.
  3. Set the rejection-sampling cadence and sample size in section 5.3 from your own volume and risk tolerance, not a generic default.
  4. Point the cross-references in sections 2, 5.2, and 5.5 to your real case-processing, signal-management, and deviation procedures.
  5. Confirm every regulation in section 7 against the current published version before issue.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.