Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
Protocol Plug-and-play starting point AI & Automation

Protocol: AI/ML Validation for Pharmacovigilance Case-Processing Models

A plug-and-play validation protocol for an AI or machine learning model used in pharmacovigilance, such as MedDRA auto-coding or signal-triage augmentation: risk-pattern classification, safety-relevant performance criteria, verbatim and negation test cases, MedDRA version pinning, rejection sampling for screening models, and a filled specimen with a worked confusion matrix.

Document type: Protocol

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use validation protocol for an AI or machine learning model used in a pharmacovigilance process, for example MedDRA auto-coding, intake triage, literature screening, or signal-detection augmentation. It builds on the general method in Protocol: AI/ML system validation and adds the test cases that are specific to drug safety: recall on the safety-relevant class, verbatim preservation, negation and uncertainty handling, MedDRA version pinning, and rejection sampling for a model that screens what a human sees. Replace every <<FILL: ...>> placeholder with your own specifics, set your document numbers and dates, attach your requirements and traceability, and route it through your normal document control, review, and approval. A worked filled specimen follows the template. The AI-specific regulatory framework for pharmacovigilance is still forming, so verify each cited reference against the current source and state plainly in the protocol where you are reasoning by analogy because no guidance fits cleanly. This content is educational reference, not legal or regulatory advice.

Document control header

FieldEntry
Document titleAI/ML Validation Protocol for <<FILL: MODEL NAME>>, a pharmacovigilance <<FILL: use case, e.g. MedDRA coding suggestion / intake triage / signal-detection augmentation>> model
Document number<<FILL: PRT-ID, e.g. PRT-PV-AI-002>>
Version<<FILL: version, e.g. 1.0>>
Effective date<<FILL: effective date>>
Supersedes<<FILL: prior version or "New">>
Document owner<<FILL: role, e.g. PV Systems Owner>>
Model identifier and version<<FILL: model name, version or hash, framework>>
MedDRA version pinned (if coding-relevant)<<FILL: version, e.g. 27.1>>
Linked URS / risk pattern assessment<<FILL: URS-ID; AI risk pattern worksheet ID>>
Linked report<<FILL: validation summary report number to be issued>>

1. Objective

This protocol demonstrates that <<FILL: MODEL NAME>>, model version <<FILL: model version>>, meets its approved performance specification for its intended use in the pharmacovigilance process named in section 2, behaves acceptably on negated and uncertain safety language, preserves the reporter’s verbatim where the use case touches case data, remains tied to a specific MedDRA version where it codes, can be shown to remain in a validated state over time, and operates under a defined and recorded human-oversight control appropriate to its risk pattern. The objective is to prove a performance distribution against a threshold sized to the consequence of a missed case or a missed signal, not to prove a single fixed input-output table once and freeze it.

2. Scope

This protocol covers the trained model instance named in the header, the safety data flowing into and out of it, the human-oversight step over its outputs, and the audit trail that records both the model output and the human decision. It applies to <<FILL: the specific PV process, e.g. inbound case intake triage / MedDRA coding of the verbatim / literature screening / signal-cluster review>> at <<FILL: site / business unit>>.

In scope: model performance against the specification on the safety-relevant class, verbatim handling, negation and uncertainty behavior, MedDRA version dependency, boundary and adversarial behavior, drift detection and the response to it, the human-oversight workflow, the audit trail and electronic-record controls, and traceability from the user requirements to the test evidence. For a screening model (risk pattern B, defined in section 4), rejection sampling of what the model did not surface to a human is explicitly in scope.

Out of scope: qualification of the underlying safety database or platform infrastructure (covered by <<FILL: infrastructure qualification / platform validation ID>>); the case-processing procedure itself, governed by <<FILL: SOP-ID for ICSR intake>> or <<FILL: SOP-ID for signal management>>; and periodic aggregate reporting content, which is out of scope unless the model drafts aggregate report sections, in which case <<FILL: reference the aggregate-report use case addendum>>.

3. Responsibilities

RoleResponsibility
Validation lead / authorAuthors this protocol, derives test cases from the URS performance specification, maintains traceability, manages deviations, writes the summary report.
PV / Safety System OwnerOwns the intended-use statement and risk pattern; confirms the test scenarios reflect real case flow; accepts the result for the process.
Data science / ML engineeringProvides the model, the training and test data records, the locked held-out test set, and the technical evidence for drift; supports execution.
Data steward / qualified safety labelersProvide and qualify ground-truth labels (validity, seriousness, MedDRA term) for test data; record inter-rater agreement.
Safety physician / PV scientistConfirms the clinical correctness of ground-truth labels for signal and narrative test cases; reviews explainability outputs for sense.
Trained case processors / coders (end-user reviewers)Execute the human-oversight test cases as the real role; record decisions and evidence.
Quality AssuranceReviews and approves the protocol, the acceptance criteria, the deviation dispositions, and the summary report; owns release for GxP use.
QPPV or delegateInformed of the risk pattern and residual risk for any model touching signal or reporting-clock decisions.

4. Intended use and risk pattern

State the intended use in one sentence naming the model output, the action it triggers, and the accountable role. If you cannot write that sentence cleanly, the intended use is not yet defined and validation cannot be sized correctly.

FieldEntry
Intended-use statement<<FILL: one sentence: the model output, the action it triggers, the accountable role>>
Risk pattern<<FILL: A, human-confirmed assistance / B, model-gated screening / C, autonomous or near-autonomous action>>
Decision the output drives<<FILL: what happens because of the output, e.g. sets the coded Preferred Term / routes the case to a processor / discards a "not relevant" message>>
Accountable role for the consequence<<FILL: role that owns the final decision>>
Rationale for the risk pattern<<FILL: ICH Q9(R1)-based rationale on file>>
Consequence of a wrong output<<FILL: a missed case, a mis-coded term that fragments signal statistics, a missed reportable article, a distorted signal recommendation; state which failure is worse and why>>

Pattern A is the lowest-risk pattern: a qualified person reviews every output before it has any effect. Pattern B is a model that decides what a human sees, so the dangerous failure is the false negative that never reaches a reviewer; this protocol adds rejection sampling (section 9.4) specifically for this pattern. Pattern C, an output that drives a regulatory-relevant action with no per-item human confirmation, requires the most evidence and a documented rationale for why a human is not in the loop; most PV deployments should not be validated as Pattern C without that rationale on file and approved by QA and the QPPV or delegate.

5. Data and model description

This section is the reproducibility record for what was validated. Without it, an investigator cannot rebuild the model and an inspector cannot trace the chain from the reporter’s original words to the coded, structured record.

5.1 Model description

FieldEntry
Model type / architecture<<FILL: e.g. text classifier, sequence-labeling coder, retrieval-augmented generative narrative assistant>>
Framework and version<<FILL: library and version, or vendor platform and version>>
Model version or hash<<FILL: immutable identifier for the exact model under test>>
Inputs<<FILL: verbatim text, structured case fields, literature abstract, source language>>
Outputs<<FILL: proposed Preferred Term, validity/seriousness classification, relevance flag, duplicate score, drafted narrative segment>>
Determinism<<FILL: deterministic at inference? any stochastic element, e.g. sampling temperature for a generative component>>
For a vendor/API base model or embedded platform feature<<FILL: base model and version pinned; vendor change behavior; whether this is a feature embedded in the safety database platform>>

5.2 Training, tuning, and test data

FieldEntry
Source system and extract date<<FILL: was the training data drawn from validated safety records under controls when created>>
Record count and date range<<FILL>>
Representativeness across the long tail<<FILL: products, sources (call center, portal, literature, partner, social media), languages, report types, and the rare serious events the model must not miss; state known gaps>>
Labeling SOP and labelers<<FILL: labeling procedure ID; qualified safety labelers and their qualification>>
Inter-rater agreement<<FILL: measured agreement for subjective labels (validity, seriousness, coded term); how disagreements were resolved>>
Class balance / base rate of the safety-relevant class<<FILL: e.g. 5% of messages are a possible adverse event>>
MedDRA version of the training and test data<<FILL: version; confirm this matches the version pinned in the header>>
Verbatim handling in the dataset<<FILL: confirm the reporter's original words were preserved and are traceable for every training and test record>>
Split method<<FILL: train / tune / test split; test set time-separated where feasible>>
Locked test set<<FILL: version hash; confirmation the model never saw it during training or tuning>>
Dataset version control<<FILL: version or hash of the frozen datasets; where retained>>

6. Acceptance criteria for a pharmacovigilance model

Acceptance criteria are performance metrics against thresholds, set before training and recorded in the URS, not after the model is measured. In pharmacovigilance the dangerous failure is almost always the miss, so recall on the safety-relevant class is usually the controlling metric; state explicitly if a different metric governs and why.

Acceptance criterionMetricThresholdJustification (consequence of error)Test population
AC-1 Primary safety performance<<FILL: recall on the safety-relevant class, e.g. possible-AE, reportable article, critical signal cluster>><<FILL: e.g. recall >= 0.94>><<FILL: a missed case or article is the dangerous failure>><<FILL: locked held-out test set, version>>
AC-2 Secondary performance<<FILL: precision on the same class>><<FILL: e.g. precision >= 0.70>><<FILL: false positives cost reviewer time>><<FILL: same test set>>
AC-3 Confidence calibration<<FILL: calibration measure>><<FILL: stated confidence matches observed accuracy within X>><<FILL: workflow routes by confidence threshold>><<FILL: test set>>
AC-4 Statistical confidence<<FILL: confidence interval on recall>><<FILL: lower CI bound still meets AC-1>><<FILL: the safety-relevant class is rare, so a point estimate is fragile>><<FILL: test set>>
AC-5 Verbatim preservationThe reporter’s original words are retained unaltered and separately from the coded output for every processed case100 percent<<FILL: the verbatim is the "Original" leg of ALCOA+ and the ground truth for recoding>><<FILL: sample of processed cases>>
AC-6 Negation and uncertainty handlingCorrect polarity on a fixed paired test set (see section 9.3)<<FILL: e.g. 100% correct on the fixed pair set>><<FILL: a flipped negation turns a non-event into a missed or false signal>><<FILL: negation/uncertainty test set>>
AC-7 MedDRA version dependency (coding models only)Coded output is produced and versioned against the pinned MedDRA release named in the header<<FILL: 100% of coded terms carry the version>><<FILL: recoding to a later version must be traceable to the version in force at the time>><<FILL>>
AC-8 Rejection-sampling false-negative rate (risk pattern B only)Measured false-negative rate among a re-reviewed sample of the model’s rejections<<FILL: e.g. no more than X% of a sized sample of rejections is a true miss>><<FILL: a screening model's dangerous failure is invisible unless sampled>><<FILL: periodic sample of rejected items>>
AC-9 Drift detectionDrift signal triggers the defined response<<FILL: detection within the defined window>><<FILL: the validated state can lapse silently>><<FILL: simulated shifted input>>
AC-10 Human oversightReview step performed, recorded, meaningful<<FILL: 100% of in-scope outputs reviewed and recorded per role, or a sized rejection sample for Pattern B>><<FILL: human judgment is the control over a wrong output>><<FILL: live or staged workflow>>
AC-11 Audit trail and recordsOutput, model version, MedDRA version, reviewer decision captured<<FILL: complete, attributable, time-stamped, cannot be disabled by ordinary users>><<FILL: a decision must be reconstructable later>><<FILL: system audit trail>>

Never accept raw accuracy on an imbalanced dataset. Where the safety-relevant class is a small minority, a model that predicts the majority class every time scores misleadingly high accuracy while missing everything that matters. Report recall, precision, and, where positives are rare, a confidence interval.

7. Pre-execution requirements

Execution does not start until all of the following are confirmed and recorded in section 13.

ItemRequirement
Approved URS with performance specURS <<FILL: URS-ID>> is approved, under change control, and contains the acceptance criteria in section 6.
Intended use and risk pattern approvedSection 4 is complete and approved by QA (and QPPV or delegate for Pattern C or any signal-touching model).
Platform / infrastructure qualified<<FILL: qualification IDs>> complete; open items assessed as not blocking.
Locked test set in placeThe held-out test set is version-controlled and confirmed unseen by the model during training or tuning.
Ground truth availableTest data is labeled by qualified safety staff; inter-rater agreement recorded.
Negation/uncertainty test set builtA fixed, versioned paired-sentence set is available (section 9.3).
Model and MedDRA version frozen for the test periodThe model version and, where relevant, the pinned MedDRA version under test are frozen; any change triggers section 11.
Reviewer trainingReviewers are trained on the model and on its known weaknesses (negation errors, weak spots on rare events, calibration limits); training records referenced.

8. Test approach

Depth of testing follows the risk pattern and the consequence of error. Each test case traces to one or more acceptance criteria in section 6 and to the URS requirement it satisfies; record the mapping in the traceability matrix. Performance is reported only on the locked held-out test set; a metric computed on training data is not validation evidence.

9. Test cases

9.1 TC-PERF: Performance against the specification

  1. Run the frozen model against the entire locked test set; do not retrain, tune, or adjust thresholds during the run.
  2. Generate the confusion matrix and compute each metric in AC-1 through AC-4.
  3. Where positives are rare, compute and record the confidence interval per AC-4.
  4. Acceptance: every metric meets or exceeds its threshold, and the lower confidence bound still meets the primary threshold.

9.2 TC-CAL: Confidence calibration

  1. Bin predictions by stated confidence.
  2. For each bin, compute the observed accuracy and compare to the stated confidence.
  3. Acceptance: stated confidence matches observed accuracy within the AC-3 tolerance, sound for any workflow that routes by confidence threshold.

9.3 TC-NEGATION: Negation and uncertainty handling

  1. Assemble a fixed, versioned set of paired sentences differing only in polarity or hedging (for example “denies chest pain” against “reports chest pain,” “no evidence of rash” against “rash noted”), plus real historical verbatims already known to contain a negation or a hedge.
  2. Run the model against the set.
  3. Acceptance: every pair resolves to the correct polarity per AC-6. Record any failure individually; a model that is otherwise accurate but fails negation systematically is not acceptable for release, because the failure lands exactly on the safety-relevant direction.

9.4 TC-VERBATIM: Verbatim preservation

  1. Process a sample of test cases end to end through the model and any downstream storage step it touches.
  2. Confirm the reporter’s original text is stored unaltered and separately from the coded or classified output, and that recoding does not overwrite it.
  3. Acceptance: verbatim is preserved unaltered for 100 percent of the sample per AC-5.

9.5 TC-MEDDRA-VERSION: MedDRA version dependency (coding models only)

  1. Confirm every coded output in the test run carries the MedDRA version pinned in the header.
  2. Simulate or apply a MedDRA version change on a small subset and confirm the model output and the change are both captured as distinct, dated entries rather than an overwrite.
  3. Acceptance: coded output is version-traceable per AC-7.

9.6 TC-REJECT-SAMPLE: Rejection sampling (risk pattern B only)

  1. Draw a representative, sized sample of items the model did not surface to a human (dropped as “not relevant,” discarded as a likely duplicate, or filtered from literature screening).
  2. Have a qualified reviewer independently assess each sampled rejection against ground truth.
  3. Compute the false-negative rate within the sample.
  4. Acceptance: the measured false-negative rate meets AC-8. This test case has no equivalent in a generic AI validation protocol; a Pattern B model without it is not validated, it is trusted, and trust is not evidence.

9.7 TC-BOUND: Boundary and out-of-distribution behavior

  1. Submit inputs at the edge of the expected range and inputs from outside the training distribution (a new product, an unfamiliar reporting channel, a foreign-language verbatim if translation is in scope).
  2. Observe whether the model flags low confidence or routes to human review rather than returning a confident wrong answer silently.
  3. Acceptance: out-of-distribution and boundary inputs produce a defined safe behavior, not a silent high-confidence error.

9.8 TC-HUMAN: Human oversight in operation

  1. Run representative outputs through the live or staged workflow with the real reviewer role.
  2. Confirm the reviewer sees the model output, its confidence or rationale, and the model version, and can confirm or override.
  3. Confirm the reviewer’s decision, the output reviewed, and the model version are recorded in a GxP record.
  4. Monitor the rate at which suggestions are accepted unmodified as a signal for automation bias; treat an unusually high rate as something to investigate.
  5. Acceptance: every in-scope output is reviewed and recorded per AC-10, and the review is defined and meaningful.

9.9 TC-AUDIT: Audit trail and electronic records

  1. Generate model outputs and corresponding human decisions, then export the audit trail.
  2. Confirm the audit trail captures the model output, the model version, the MedDRA version where relevant, the reviewer identity, the decision, and a time stamp.
  3. Acceptance: the audit trail and record controls meet AC-11.

9.10 Test case execution record (per case)

FieldEntry
Test case ID and title<<FILL>>
Traces to AC / URS<<FILL>>
Steps<<FILL>>
Expected result<<FILL>>
Actual result<<FILL>>
Pass / Fail<<FILL>>
Objective evidence reference<<FILL>>
Tester (initials, date)<<FILL>>

10. Acceptance criteria for the protocol as a whole

  • The intended-use statement and risk pattern (A, B, or C) are written and approved, with the ICH Q9(R1)-based rationale on file.
  • The performance specification was written in the URS before training, with the safety-relevant metric, its threshold, and a justification tied to the consequence of the miss.
  • Performance (AC-1 to AC-4) was reported on the locked, version-controlled test set the model never saw, and every metric meets its threshold.
  • Verbatim preservation, negation and uncertainty handling, and MedDRA version dependency (AC-5 to AC-7, where applicable) meet their criteria.
  • For a risk pattern B model, rejection sampling (AC-8) is complete and meets its criterion.
  • Drift detection, human oversight, and the audit trail (AC-9 to AC-11) meet their criteria.
  • Every open deviation is closed or assessed as not blocking, with the rationale recorded.
  • Traceability runs from intended use to requirements to test evidence, with rationale recorded wherever guidance was silent.

11. Deviation handling

  1. Record any actual result that does not meet an acceptance criterion contemporaneously with the test case ID and the objective evidence.
  2. Classify by impact: a deviation on AC-1, AC-5, AC-6, AC-7, or AC-8 (the safety-critical criteria) blocks release until resolved; a deviation on a lower-risk criterion may be dispositioned with QA-approved rationale.
  3. Investigate root cause. For a performance shortfall, determine whether the cause is the data, the labels, the metric choice, the threshold, or the model itself; do not move the threshold to make the result pass.
  4. Define the correction and re-test the affected cases.
  5. QA reviews and approves the disposition of every deviation. Route to <<FILL: SOP-ID for deviations>> where the impact warrants a formal investigation.

12. Summary and conclusion

FieldEntry
Model and version tested<<FILL>>
Test set version<<FILL>>
AC-1 to AC-11 result (met / not met)<<FILL: per criterion, with the measured value>>
Deviations raised / open<<FILL>>
Areas validated by analogy (guidance silent)<<FILL>>
Monitoring plan reference<<FILL>>
Predetermined change control plan reference<<FILL>>
Release recommendation<<FILL>>
Validation lead (name, signature, date)<<FILL>>
QA approval (name, signature, date)<<FILL>>

13. Pre-execution and execution confirmation

ItemConfirmed byDate
Pre-execution requirements (section 7) met<<FILL>><<FILL>>
Model and MedDRA version frozen<<FILL>><<FILL>>
Execution started<<FILL>><<FILL>>
Execution completed<<FILL>><<FILL>>

14. References

ICH E2B(R3), Data Elements for Transmission of Individual Case Safety Reports; ICH E2D(R1), Post-Approval Safety Data. EU GVP Module VI (case management) and Module IX (signal management). FDA draft guidance, “Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products” (issued 6 January 2025), risk-based credibility framework; draft, confirm current status. FDA and EMA, “Guiding Principles of Good AI Practice in Drug Development” (published jointly 14 January 2026), foundational, non-binding principles spanning the medicine lifecycle including post-marketing. EMA, “Reflection paper on the use of artificial intelligence in the lifecycle of medicines” (adopted September 2024), non-binding. GAMP 5 Second Edition (ISPE, 2022), risk-based CSV with material on machine learning. ICH Q9(R1), Quality Risk Management, basis for sizing validation effort to risk. 21 CFR Part 11 and EU GMP Annex 11, electronic records, signatures, and audit trails. 21 CFR 314.80 / 600.80 (postmarketing adverse experience reporting) and 21 CFR 312.32 (IND safety reporting), as applicable.

Several references above are draft or foundational rather than final and binding. Confirm the current version, status, and dates of each before issue, and state in the protocol where a reference is draft rather than settled.

15. Revision history

VersionDateAuthorSummary of change
<<FILL: 1.0>><<FILL: date>><<FILL: author>>Initial issue.

16. Approvals

RoleNameSignatureDate
Author (Validation lead)<<FILL>>
Reviewer (Data Science)<<FILL>>
Reviewer (PV / Safety System Owner)<<FILL>>
Approver (Quality Head)<<FILL>>

Filled specimen

The following shows the protocol completed for an example MedDRA coding-suggestion model with mandatory human confirmation (risk pattern A). The company, model, and numbers are illustrative; replace them with your own.

Intended use and risk pattern (filled)

FieldEntry
Intended-use statementThe model reads the case verbatim and proposes a MedDRA Preferred Term; a trained coder confirms or overrides the term before the case record is finalized; the coder owns the final coded term.
Risk patternA, human-confirmed assistance
Decision the output drivesThe proposed term populates the coding field for coder review; no term is saved without coder confirmation
Accountable roleTrained safety coder
RationaleEvery output is reviewed before it has any effect on the case record; residual risk is that coders rubber-stamp a usually-correct suggestion, mitigated by the override-rate monitoring in section 9.8
Consequence of a wrong outputA wrong suggestion that a coder accepts without checking corrupts the coded term and, in aggregate, distorts signal statistics that depend on coding consistency

Acceptance criteria (filled, primary lines)

Acceptance criterionMetricThresholdTest population
AC-1 Primary safety performanceExact-match agreement with the coder-confirmed Preferred Term>= 0.85Locked test set v2, 2,000 verbatims, current MedDRA version 27.1
AC-6 Negation and uncertaintyCorrect polarity on paired set100% of 60 pairsFixed negation/uncertainty set NEG-PV-001
AC-5 Verbatim preservationUnaltered, separate storage100% of 200-case sampleSample of processed cases

Performance result (filled, from TC-PERF)

On locked test set v2 (2,000 verbatims), the model’s top suggestion exactly matched the coder-confirmed term in 1,720 of 2,000 cases, an agreement rate of 0.86, meeting AC-1. On TC-NEGATION, all 60 paired sentences resolved to correct polarity. On TC-VERBATIM, all 200 sampled cases retained the original verbatim unaltered and separately from the coded field.

Deviation (filled)

FieldEntry
Deviation IDDEV-PV-AI-2026-0003, raised against TC-PERF
FindingAgreement fell to 0.71 specifically on verbatims describing injection-site reactions, a category underrepresented in training (38 of 2,000 test cases)
Root causeTraining data drawn primarily from oral-dose products; the coding model’s site-of-administration vocabulary was thin
CorrectionTraining set expanded with 400 additional injection-site verbatims from the historical archive, relabeled by two qualified coders with agreement recorded; model retrained; TC-PERF re-run
Re-test resultOverall agreement 0.87; injection-site subset agreement 0.83, within the tolerance QA accepted given the small subgroup size
QA dispositionClosed; not blocking after re-test. Approved by R. Gomez, 09 August 2026

In this example the subgroup weakness was found by looking past the aggregate number, not by trusting it. That is exactly what section 6’s insistence on representativeness across the long tail is for.

Common inspection findings this protocol prevents

  • A coding model validated only on aggregate agreement, hiding a systematic weakness on an underrepresented product class or event type.
  • No verbatim-preservation test, so recoding to a later MedDRA version silently overwrites the reporter’s original words.
  • No negation testing, so a model that reliably misreads “denies X” as “reports X” reaches production undetected.
  • A screening model with no rejection-sampling evidence, so the false-negative rate the model produces in production is unknown.
  • Coded output with no MedDRA version stamp, so a later dictionary change cannot be reconciled against what was coded and when.
  • Acceptance criteria written after the model’s performance was known.

How to adapt this protocol

  1. Set your document number, model identifier, MedDRA version, and effective date in the header.
  2. Write the intended-use sentence and assign the risk pattern (A, B, or C) in section 4 before anything downstream is sized.
  3. Fill the data and model description in section 5 completely, including the MedDRA version of the training and test data.
  4. Set the acceptance criteria in section 6 before training, and keep AC-5 through AC-8 even if your first instinct is to skip them; they are the PV-specific evidence a generic AI protocol does not ask for.
  5. Build the negation/uncertainty test set (9.3) from your own historical verbatims, not a generic sentence list.
  6. If the model is risk pattern B, do not release without executing TC-REJECT-SAMPLE and sizing the sample to detect a meaningful false-negative rate.
  7. Confirm every reference in section 14 against the current published version and status before issue.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.