Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
Checklist Plug-and-play starting point AI & Automation

Checklist: Good Machine Learning Practice (GMLP) Conformance and Evidence

A plug-and-play checklist that maps the ten GMLP guiding principles to the real evidence objects a reviewer expects, with a pass/partial/gap column, an evidence reference for each, and a filled specimen. GMLP is a lens for completeness, not a certificate.

Document type: Checklist

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use conformance checklist for Good Machine Learning Practice (GMLP). GMLP is not something you certify against; it is a set of habits a credible AI program demonstrates, so this checklist points to where each principle is satisfied in your real records rather than asking you to write a document called “GMLP compliance.” Each principle below is described in original wording; confirm the current guiding-principles text against the source before relying on it. Replace every <<FILL: ...>> placeholder. A worked filled specimen follows. This is general guidance to adapt, not legal or regulatory advice.

Control header

FieldEntry
Checklist number<<FILL: FRM-ID>>
Product / model<<FILL: AI-enabled device or SaMD name>>
Intended use<<FILL: clinical task, population, workflow>>
Model type<<FILL: locked / adaptive; architecture family>>
Reviewer / date<<FILL>>
QA / Regulatory sign-off / date<<FILL>>

How to use

For each principle, name the evidence object that satisfies it and its document reference, then mark the status. “Pass” means the evidence exists and is under document control. “Partial” and “Gap” become actions. GMLP is a completeness lens: a Gap here usually maps to a review or inspection finding.

The ten principles, evidence, and status

The principles are the ten jointly published by FDA, Health Canada, and the MHRA (October 2021), described here in original wording for working use.

#Principle (described)Evidence object to point toReferenceStatus (Pass / Partial / Gap)
1Every discipline is engaged from the start and stays engaged, not consulted at the endDesign-review attendance, RACI, risk file showing clinical, data science, software, quality, and human factors involved throughout<<FILL>><<FILL>>
2The model rests on sound software engineering and securityVersion control records, secure-development evidence, data-integrity controls, cybersecurity threat model<<FILL>><<FILL>>
3The data looks like the people who will actually be exposed to the productData card / data sheet with source, provenance, and demographic composition across age, sex, race, ethnicity, severity, sites, equipment<<FILL>><<FILL>>
4The data the model learns from is kept separate from the data used to judge itData-partitioning record showing no leakage between training, tuning, and locked test sets<<FILL>><<FILL>>
5Ground truth is defined with the strongest method availableGround-truth definition and adjudication method, reference-standard rationale<<FILL>><<FILL>>
6The model is fitted to the data on hand and to the intended useModel-selection rationale, mapping of model outputs to clinical decisions<<FILL>><<FILL>>
7What is measured is the clinician-plus-tool, not the algorithm on its ownHuman-factors study, reader study of clinicians with vs without the tool<<FILL>><<FILL>>
8The model is tested under conditions that reflect real clinical useValidation report on held-out data, realistic sites, workflow, and subgroup breakdown<<FILL>><<FILL>>
9Users get the information they need to use it wellLabeling and transparency artifacts: intended use, performance with uncertainty, validated population, limitations, active version<<FILL>><<FILL>>
10Once the product ships, someone keeps checking whether it still performs, and has a plan for what a future retrain would need before it goes liveReal-world performance monitoring plan and records: performance vs baseline, input drift, override rate, defined response<<FILL>><<FILL>>

Actions from gaps

Principle #Gap / partial detailActionOwnerDue
<<FILL>><<FILL>><<FILL>><<FILL>><<FILL>>

References

Good Machine Learning Practice for Medical Device Development: Guiding Principles, jointly published by FDA, Health Canada, and the MHRA (October 2021), ten principles. FDA final guidance, Marketing Submission Recommendations for a Predetermined Change Control Plan for AI-Enabled Device Software Functions (final December 2024, reissued August 2025). Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles (jointly published June 2024).

Describe these principles in your own wording in any submission; confirm the current text before relying on it.


Filled specimen (excerpt)

Illustrative rows for a locked chest X-ray triage SaMD.

#Principle (described)Evidence objectReferenceStatus
3Data represents the intended populationData card v2.1: 41,000 studies, 9 sites, 3 scanner vendors, age/sex/site breakdown; one vendor under 10% of data, flaggedDATA-CARD-CXR-021Partial (vendor B under-represented, monitored post-market)
4Training and test independencePartition log: patient-level split, no cross-site leakage, locked test set touched onceDS-SPLIT-014Pass
7Human-AI team performanceReader study, 12 radiologists, with vs without tool, prioritization time and miss rateHF-STUDY-007Pass
10Post-market monitoringMonitoring plan live from launch: monthly performance vs baseline, weekly input-drift check, override-rate dashboard, defined pause-and-escalate responsePMS-PLAN-003Pass

The Partial on principle 3 is honest: the under-represented scanner vendor is named, flagged as a transparency limitation, and made a monitoring target rather than buried under the aggregate number.

Common findings this checklist prevents

  • A model validated on one or two sites, then claimed for general use, with no representativeness evidence.
  • No subgroup analysis, so a performance gap for an under-represented group is invisible.
  • Test-set leakage inflating the headline number.
  • Standalone algorithm metrics presented as clinical performance with no human-AI study.
  • A “GMLP compliance” document that asserts conformance but points to no real evidence object.

How to adapt this checklist

  1. Enter your model, intended use, and model type in the header.
  2. For each principle, name the actual evidence object and its controlled reference, then mark the status honestly.
  3. Turn every Partial and Gap into a dated action.
  4. Reuse the completed checklist as the completeness map for a design review or an inspection.
  5. Confirm the current guiding-principles text against the source before relying on it.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.