Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
Checklist Plug-and-play starting point AI & Automation

Checklist: Generative AI / LLM Qualification Release Readiness

A plug-and-play pre-release go/no-go checklist for a GxP generative AI system: intended use, probabilistic acceptance criteria, golden dataset, scoring and any LLM judge, prompt and version control, guardrails, and production monitoring, each with pass/fail/NA and a signed disposition, with a filled specimen.

Document type: Checklist

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use release-readiness checklist. Run it as the go/no-go gate before a generative AI system produces or drives any GxP record. It pulls the qualification method into a single signed decision, so a reviewer and QA can see at a glance whether every load-bearing control is in place. Replace each <<FILL: ...>> placeholder with your own specifics. A filled specimen follows. Mark each item Pass, Fail, or N/A with evidence; any Fail on a required item blocks release.

Document control header

FieldEntry
Checklist titleGenerative AI Qualification Release Readiness
System / use case<<FILL: name>>
Version qualified<<FILL: model + prompt + retrieval config version>>
Reviewer<<FILL: role>>
Date<<FILL: date>>

1. Intended use and risk pattern

#ItemPass/Fail/NAEvidence
1.1An intended-use statement names the generative output, the action it triggers, and the accountable role<<FILL>>
1.2The use case is assigned to a generative risk pattern (drafting with human authorship, retrieval QA, structured extraction, or autonomous action), with an ICH Q9(R1) rationale<<FILL>>
1.3If the output drives action with no per-output human review, the highest-risk controls are in place or release is refused<<FILL>>

2. Acceptance criteria

#ItemPass/Fail/NAEvidence
2.1More than one metric is defined (groundedness, hallucination rate, completeness, consistency, abstention, accuracy where checkable), not a single accuracy number<<FILL>>
2.2Each metric has a numeric threshold justified against the consequence of error<<FILL>>
2.3The thresholds were recorded in the URS before any evaluation was run<<FILL>>

3. Golden dataset

#ItemPass/Fail/NAEvidence
3.1The dataset is representative and includes edge, unanswerable, and adversarial cases<<FILL>>
3.2Expected behavior was set by qualified SMEs under a labeling SOP with recorded inter-rater agreement<<FILL>>
3.3The set is sized to support the thresholds with a stated confidence interval<<FILL>>
3.4Development and locked-evaluation splits are separated; the locked split was not used for tuning<<FILL>>
3.5The dataset is versioned and hashed<<FILL>>

4. Scoring and any LLM judge

#ItemPass/Fail/NAEvidence
4.1The scoring method for each metric is defined<<FILL>>
4.2Any LLM judge is qualified: measured agreement with human experts above a stated threshold, per category<<FILL>>
4.3The human-scoring rubric and rater qualifications are on file<<FILL>>

5. Prompt and version control

#ItemPass/Fail/NAEvidence
5.1Prompt, retrieval config, guardrails, model version, and sampling parameters are under version control<<FILL>>
5.2The model version is pinned, not a “latest” alias<<FILL>>
5.3Each qualification result references the exact configuration version it was produced under<<FILL>>
5.4GxP outputs are traceable to the configuration that generated them<<FILL>>

6. Guardrails

#ItemPass/Fail/NAEvidence
6.1Each guardrail is specified as a requirement and verified with scripted test cases, including negative and adversarial inputs<<FILL>>
6.2Outputs are grounded in and traceable to retrieved sources<<FILL>>
6.3Structured outputs are schema-validated; injection and out-of-scope inputs are filtered<<FILL>>
6.4A human review gate exists for any output that becomes or drives a GxP record<<FILL>>

7. Production monitoring

#ItemPass/Fail/NAEvidence
7.1A monitoring plan is live from day one with defined triggers (cadence, override-rate threshold, input drift, vendor version change)<<FILL>>
7.2Re-evaluation on the locked golden dataset is scheduled and triggered by any configuration change<<FILL>>
7.3A defined response exists when a monitor trips (who is notified, pause or route to fuller review, how it is investigated)<<FILL>>

8. Records and integrity

#ItemPass/Fail/NAEvidence
8.1Any output that becomes part of a GxP record is treated under Part 11 and ALCOA+, not as informal<<FILL>>
8.2Any embedded or vendor generative feature that touches a GxP process is inventoried and assessed, not hidden<<FILL>>

Disposition

FieldEntry
Any required item Failed?Yes / No (Yes blocks release)
Open items with justification<<FILL: list or none>>
DecisionRelease / Do not release / Release with restriction
Reviewer signature and date<<FILL>>
QA approval signature and date<<FILL>>

References

21 CFR Part 11 and ALCOA+ for record integrity. ICH Q9(R1) (Quality Risk Management) for risk-based sizing of the qualification. FDA guidance on Computer Software Assurance for Production and Quality Management System Software (current version). GAMP 5 (second edition) risk-based principles; described in original wording, not reproduced.

Confirm the current version of each reference before you rely on it.

Filled specimen (extract)

Illustrative; replace with your own. Retrieval-QA assistant answering specification questions.

#ItemResultEvidence
2.1Multiple metrics definedPassURS-AI-07: groundedness, hallucination, completeness, consistency, abstention, numeric accuracy
3.4Splits separated, locked split not tunedPassDataset SPEC-AI-03: 120 dev / 280 locked, hashes on file
4.2LLM judge qualifiedPass (partial)Judge protocol PRO-AI-11: qualified for groundedness/completeness on answerable and unanswerable items; adversarial items human-scored
6.4Human review gate presentPassAnswer is reference only; controlled document remains the record; no disposition authorized by the assistant
7.1Monitoring live from day onePassMonitoring plan MON-AI-05: weekly override rate, monthly re-eval, model-version watch

Disposition in the specimen: no required item Failed, the judge limitation on adversarial items is documented with human-scoring fallback, so the decision was Release with the monitoring plan active. The value of the checklist is that it forces the one question that sinks weak generative deployments to be answered before go-live: is there a real human gate, a locked golden dataset, a pinned model, and a live monitor, or just a single accuracy number.

Common inspection findings this checklist prevents

  • A generative system released on a single accuracy number with no golden dataset behind it.
  • A “latest” model alias in production, so the running system is not the qualified one.
  • No monitoring, so a silent vendor model change degrades the system unnoticed.
  • An embedded vendor generative feature writing into records, never assessed as a generative system at all.

How to adapt this checklist

  1. Set the system, qualified version, and reviewer in the header.
  2. Mark items that genuinely do not apply as N/A with a reason; do not leave them blank.
  3. Treat any Fail on a required item as a release blocker, and record any release-with-restriction decision and its basis.
  4. Run the checklist again after any configuration change that triggers re-evaluation, not only at first release.
  5. Confirm every regulation in the references against the current published version before you rely on it. This checklist is educational guidance for adapting your own qualification and release process; it is not legal, regulatory, or professional advice, and it does not itself certify or guarantee compliance with any regulation or guidance it references.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.