Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
Record Plug-and-play starting point AI & Automation

Record: AI/ML Training Data Integrity and Dataset Version

A plug-and-play record for the integrity of a GxP model's training data: source and lineage, representativeness, labeling quality and inter-rater agreement, class balance, and the frozen dataset version, so the model can be reproduced and investigated, with a worked specimen.

Document type: Record

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use record for the integrity of a GxP AI/ML model’s training data. The integrity of a supervised model is bounded by the integrity of the data it learned from, so this record is a GxP record in its own right: it lets a future investigator reproduce the model and lets an inspector trace the chain from data to decision. Replace every <<FILL: ...>> placeholder with your own specifics. A worked filled specimen follows. Verify each cited regulation against the current source before you rely on it.

Document control header

FieldEntry
Record titleTraining Data Integrity and Dataset Version, <<FILL: MODEL / SYSTEM NAME>>
Document number<<FILL: REC-ID, e.g. AI-DATA-011>>
Version<<FILL: version>>
Effective date<<FILL: date>>
Data steward<<FILL: role>>
Dataset version / hash<<FILL: frozen dataset identifier or hash>>

1. Source and lineage

FieldEntry
Source system(s)<<FILL: system name(s)>>
Extract date / range<<FILL: extract date; data date range>>
Was the data created under GxP controls?<<FILL: yes/no; notes>>
Number of records<<FILL: count>>
Extraction method / query<<FILL: how the data was pulled; ref>>
Chain from source to frozen set<<FILL: transformations applied, each recorded>>

2. Representativeness

FieldEntry
Products / processes represented<<FILL>>
Sites / instruments represented<<FILL>>
Edge and rare cases included<<FILL: how rare/critical cases are represented>>
Known gaps vs the production population<<FILL: what the data under-represents, and the risk>>

State honestly what the data does not cover. A model trained only on routine cases will fail on the rare events that matter most, and an unstated gap is a finding.

3. Labeling quality (supervised models)

FieldEntry
Labeling SOP / procedure<<FILL: SOP-ID>>
Labelers and qualifications<<FILL: names/roles; qualification basis>>
Label definitions / rubric<<FILL: reference>>
Inter-rater agreement (for subjective labels)<<FILL: metric and value, e.g. Cohen's kappa>>
Disagreement resolution method<<FILL: how conflicts were resolved and recorded>>
Ground-truth basis<<FILL: how the "correct" label was established>>

4. Class balance and splits

FieldEntry
Class distribution (base rates)<<FILL: e.g. 2% critical, 98% non-critical>>
Imbalance handling<<FILL: metrics chosen, sampling/weighting, if any>>
Train / validation / test split method<<FILL: method; time-based split if used>>
Test set held out and locked<<FILL: yes; version ref>>
Leakage check<<FILL: confirmation no record appears in more than one split>>

5. Retraining data provenance (if applicable)

Data pulled from production to retrain a deployed model is itself GxP data and carries the full ALCOA+ expectations. Retraining is a data-integrity event, not just an engineering task.

FieldEntry
Production data used for retraining?<<FILL: yes/no>>
ALCOA+ controls on that data<<FILL: attributable, contemporaneous, original/true copy, etc.>>
Change control reference for the retrain<<FILL: ref>>

6. Acceptance criteria for this record

  • Source, extract date, record count, and lineage are stated and reproducible.
  • Representativeness is described, including honest statement of gaps.
  • For supervised models, the labeling SOP, labeler qualifications, inter-rater agreement, and disagreement resolution are recorded.
  • Class balance is known and the metrics chosen are not fooled by imbalance.
  • The train/validation/test split is defined, the test set is locked and version-controlled, and a leakage check is recorded.
  • The frozen dataset is versioned so the model can be rebuilt.
  • Any production data used for retraining carries ALCOA+ controls and a change-control reference.

7. References

21 CFR Part 11 and EU GMP Annex 11 for electronic records around the dataset. MHRA GxP Data Integrity Guidance and Definitions; PIC/S PI 041 (reference by title; describe, do not paste). ALCOA+ data-integrity principles (attributable, legible, contemporaneous, original, accurate, plus complete, consistent, enduring, available). GAMP 5 Second Edition (ISPE, 2022) for the AI/ML lifecycle context (reference by title). Track the draft EU GMP Annex 22 and Annex 11 revision (2025); confirm status before relying on either.

Confirm the current version and clause numbers of each reference before issue.

8. Revision history

VersionDateAuthorSummary of change
<<FILL: 1.0>><<FILL: date>><<FILL>>Initial record for the frozen dataset.

9. Approvals

RoleNameSignatureDate
Data steward<<FILL>>
Data Science<<FILL>>
QA<<FILL>>

Filled specimen

The following shows the record completed for an illustrative complaint-classification model. Numbers are illustrative.

FieldEntry
Source systemComplaint management system CMS-1
Extract date / rangeExtracted 15 July 2026; records dated Jan 2023 - Jun 2026
Number of records18,400 complaints
Products / sites represented6 products, 3 sites; injectable and oral solid
Known gapsCombination-product complaints under-represented (n=140); flagged as a monitoring risk
Labeling SOPSOP-QA-221, complaint category rubric v3
LabelersTwo qualified complaint analysts (J. Ruiz, P. Adeyemi)
Inter-rater agreementCohen’s kappa 0.86 on a 400-record double-labeled subset
Disagreement resolutionThird senior reviewer adjudicated; decisions logged
Class distribution7% safety-relevant, 93% non-safety
Split methodTime-based: train Jan 2023 - Dec 2025, tune Jan - Mar 2026, test Apr - Jun 2026
Test set lockedYes, dataset version CMS1-2026-07-15-frozen, hash on file

This package lets a future investigator rebuild the exact model and shows an inspector the chain from source data to the labels to the frozen set. The honest note about under-represented combination-product complaints is exactly the kind of stated gap that reads as a controlled program rather than a hidden weakness.

Common inspection findings this record prevents

  • A model that cannot be reproduced because the training dataset was never versioned.
  • Labels generated with no documented SOP or labeler qualifications.
  • No inter-rater agreement for subjective labels, so label consistency is unknown.
  • A hidden data gap that surfaces as production failures on an under-represented population.
  • Production data used to retrain with none of the ALCOA+ controls that GxP data requires.

How to adapt this record

  1. Fill the source, extract, and lineage so the dataset can be rebuilt from scratch.
  2. State representativeness honestly, including what the data under-represents.
  3. Record the labeling SOP, labeler qualifications, inter-rater agreement, and how disagreements were resolved.
  4. Define the splits, lock the test set, version the frozen dataset, and record a leakage check.
  5. If you retrain on production data, attach the ALCOA+ controls and the change-control reference.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.