This is a ready-to-use validation plan for an AI/ML system used in a GxP process. It scopes the whole validation, states the AI-specific approach where traditional CSV does not fit, and names the deliverables, roles, and acceptance up front. The execution then runs against a protocol; this plan is the governing document above it. Replace every <<FILL: ...>> placeholder with your own specifics and route it through your validation and QA approval. A worked filled specimen follows. Verify each cited regulation against the current source before you rely on it.
Document control header
| Field | Entry |
|---|---|
| Document title | Validation Plan, <<FILL: AI SYSTEM / MODEL NAME>> |
| Document number | <<FILL: VP-ID, e.g. VP-AI-004>> |
| Version | <<FILL: version>> |
| Effective date | <<FILL: date>> |
| System owner | <<FILL: role>> |
| Approved by | <<FILL: QA>> |
1. Scope and intended use
State exactly what the system does and what role the AI output plays, in one sentence naming the output, the action it triggers, and the accountable role:
<<FILL: intended-use sentence>>
In scope: <<FILL: the model, its inputs, its outputs, the workflow it sits in>>. Out of scope: <<FILL: adjacent systems validated elsewhere, e.g. the underlying LIMS/MES>>.
2. Risk class and risk basis
| Field | Entry |
|---|---|
| AI use pattern | `<<FILL: advisory/screening |
| Risk class rationale (ICH Q9(R1)) | <<FILL: how influential the output is and how serious the decision is>> |
| GAMP software category | <<FILL: Category 4 configured, or Category 5 custom; note the trained instance is bespoke>> |
| Supplier assessment reference (if vendor/API model) | <<FILL: ref; pin the model version where the vendor allows>> |
Validation effort scales to this risk class. A process-control model carries failure-mode analysis and an independent deterministic interlock; an advisory model does not.
3. Validation approach
The plan follows the familiar lifecycle with AI-shaped content inside each stage. Where existing guidance is silent, the rationale is documented so it survives inspection.
- Requirements and performance spec. Written before training, with metrics, thresholds justified by the consequence of error, and the test population. See the performance specification deliverable.
- Training data integrity. Source, lineage, representativeness, labeling quality, class balance, and a frozen, versioned dataset. See the training data integrity record.
- Model development and testing. Development documented; performance reported on a locked held-out test set the model never saw.
- Change control. A predetermined change control plan classifying anticipated changes and their required testing, including vendor-driven base-model changes.
- Performance monitoring. Live from day one, with triggers (drift, override rate, cadence) and a defined response.
- Explainability and human review. Scaled to the use pattern; the human review step defined, documented, and kept meaningful.
4. Deliverables
| Deliverable | Reference | Owner |
|---|---|---|
| Intended-use and risk classification | <<FILL>> | System Owner |
| Performance and acceptance specification (pre-training) | <<FILL>> | System Owner / Data Science |
| Training data integrity and dataset version record | <<FILL>> | Data Steward |
| Model development / testing report (held-out test set) | <<FILL>> | Data Science |
| AI/ML risk assessment | <<FILL>> | Validation / QA |
| Predetermined change control plan | <<FILL>> | System Owner + QA |
| Performance monitoring plan | <<FILL>> | System Owner + Data Science |
| Human-review procedure and training | <<FILL>> | Operations + QA |
| Validation summary report | <<FILL>> | Validation lead |
| Traceability (intended use to requirements to test evidence) | <<FILL>> | Validation lead |
5. Roles and responsibilities
| Activity | Accountable | Contributes |
|---|---|---|
| Intended use and risk class | System Owner | QA, Data Science |
| Performance spec and requirements | System Owner | QA, SMEs, Data Science |
| Training data integrity | Data Steward / SME labelers | Data Science, QA |
| Model development and testing | Data Science / ML Engineering | System Owner |
| Validation approach and protocols | Validation / CSV lead | QA, Data Science |
| Performance and release approval | QA | System Owner |
| Post-deployment monitoring | System Owner + Data Science | QA |
Quality is involved while the model is built, not only at the end. Treating model building as a pure data-science task that QA reviews last is the recurring failure.
6. Schedule
| Milestone | Target date | Dependency |
|---|---|---|
| Intended use and risk class approved | <<FILL>> | - |
| Performance spec approved (before training) | <<FILL>> | Risk class |
| Training data record frozen | <<FILL>> | Data extract |
| Model tested on held-out set | <<FILL>> | Frozen data, spec |
| Monitoring live | <<FILL>> | Deployment |
| Validation summary approved | <<FILL>> | All above |
7. Acceptance criteria for release
The system is releasable when all of the following can be evidenced from the file:
- Intended use, risk class, and GAMP category are stated and approved.
- The performance spec was written before training, and the reported metrics on a locked held-out test set meet it.
- The training data integrity record is complete and the dataset is versioned.
- A predetermined change control plan and a live monitoring plan are approved.
- The human review step is defined, documented, and reviewers are trained on the model’s weaknesses.
- For process control, deterministic safety interlocks are independently validated.
- Traceability runs from intended use to requirements to test evidence, with rationale recorded wherever guidance was silent.
8. References
FDA guidance, Computer Software Assurance for Production and Quality Management System Software (draft September 2022; final 24 September 2025; current version issued 3 February 2026). GAMP 5 Second Edition (ISPE, 2022), including its material on AI/ML (reference by title; describe, do not paste). ICH Q9(R1), Quality Risk Management. 21 CFR Part 11 and EU GMP Annex 11 for electronic records and signatures. Track the draft EU GMP Annex 22 (Artificial Intelligence, 2025), the draft Annex 11 revision (2025), the FDA January 2025 draft on AI for regulatory decision-making, and the January 2026 FDA-EMA Guiding Principles of Good AI Practice in Drug Development; all are draft or principle-level, confirm status before relying on them.
Confirm the current version and clause numbers of each reference before issue.
9. Revision history
| Version | Date | Author | Summary of change |
|---|---|---|---|
<<FILL: 1.0>> | <<FILL: date>> | <<FILL>> | Initial plan. |
10. Approvals
| Role | Name | Signature | Date |
|---|---|---|---|
| System Owner | <<FILL>> | ||
| Validation lead | <<FILL>> | ||
| QA | <<FILL>> |
Filled specimen
The following shows the plan header and approach completed for an illustrative deviation-triage model. Details are illustrative.
- Intended use: the model assigns a preliminary criticality tier to each new deviation; the tier sets the investigation timeline; a QA reviewer confirms or overrides within one business day and owns the final tier.
- Use pattern / risk class: automated classification that drives a timeline; medium-high risk because a misclassification can delay a safety-relevant investigation. Rationale filed under RA-AI-004.
- GAMP category: Category 4 platform with a Category 5 trained model inside it; effort sized to the trained model.
- Approach highlights: performance spec URS-AI-009-PERF approved before training (recall >= 0.90 on the critical tier); training data frozen as CMS1-2026-07-15; predetermined change control plan permitting quarterly retrains with a confirmatory test; monitoring live from day one on override rate and input drift; QA reviewer confirms every tier for the first 90 days, then by risk-based sampling.
This plan reads as a system under control: the sizing is justified, the spec preceded the training, the change and monitoring plans exist before go-live, and the human control is defined. That coherence is what an inspector rewards when no AI-specific standard yet fits cleanly.
Common inspection findings this plan prevents
- Validation effort mis-sized because the intended use and risk class were never pinned down.
- A validation “protocol” with no governing plan, so scope and acceptance drift.
- No change control or monitoring plan defined before go-live, so the validated state lapses silently.
- A process-influencing model validated as if advisory, with no failure-mode analysis or interlock.
- Traceability that cannot connect intended use to the test evidence that proves it.
How to adapt this plan
- Write the intended-use sentence first; if you cannot write it cleanly, the scope is not defined.
- Assign the risk class and size every downstream deliverable to it.
- List the deliverables and owners; each links to its own template.
- Set the schedule so the performance spec is approved before training and monitoring is live at go-live.
- Confirm every regulation in the references against the current published version, including the draft status of the AI-specific frameworks, before issue.