Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
SOP Plug-and-play starting point AI & Automation

SOP: AI/ML Ground Truth Labeling

A plug-and-play standard operating procedure for building trustworthy ground truth to train AI or machine learning models: writing precise category definitions with edge-case examples, labeler qualification, double-labeling a subset, measuring inter-rater agreement against a pre-set threshold, adjudicating disagreement, and versioning the label set, with a worked agreement calculation.

Document type: SOP

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use SOP. Replace every <<FILL: ...>> placeholder with your own specifics, set your document numbers and dates, and route it through your normal document control, review, and approval. A worked filled specimen, including an inter-rater agreement calculation, follows the template. Verify each cited regulation against the current source before you rely on it.

Document control header

FieldEntry
Document titleAI/ML Ground Truth Labeling
Document number<<FILL: SOP-ID, e.g. SOP-AI-014>>
Version<<FILL: version, e.g. 1.0>>
Effective date<<FILL: effective date>>
Supersedes<<FILL: prior version or "New">>
Document owner<<FILL: role, e.g. Quality Assurance / Data Science Lead>>
Applies to<<FILL: sites / departments / model types in scope>>

1. Purpose

This procedure defines how <<FILL: COMPANY NAME>> establishes ground truth, the labeled examples an AI or machine learning model learns from, so that labels used to train a GxP-relevant model reflect a defensible, consistent, qualified judgment rather than one analyst’s unrecorded habit. A label is a data point the same way a measurement is, and it carries the same expectation of being attributable, accurate, and reproducible.

2. Scope

This procedure applies whenever a person assigns a label, category, or judgment-based outcome to a record for the purpose of training, validating, or testing an AI/ML model used in a GxP-relevant workflow at the sites in the header. It covers new labeling programs and the re-use of existing structured categorization (for example a historical QMS category) as ground truth. It does not apply where the label is an objective downstream outcome (a measured result, a batch’s actual disposition) rather than a human judgment; those are governed by the source record’s own procedure.

3. Responsibilities

RoleResponsibility
Labeling program ownerWrites and maintains the category definitions, selects and qualifies labelers, and owns the labeling SOP’s fitness for purpose.
Qualified labelers (SME)Apply labels per the definitions in section 5.1; participate in double-labeling and calibration sessions.
AdjudicatorA senior reviewer, not among the original labelers of the disputed item, who resolves disagreements per section 5.5.
Data ScienceConsumes the versioned label set; reports back on any pattern of model error that suggests a label-definition problem.
Quality AssuranceReviews and approves the labeling SOP, the labeler qualification records, and the inter-rater agreement result before the label set is used to train a model.

4. Definitions

  • Ground truth: the set of labels the organization trusts enough to teach a model from, established under this procedure.
  • Inter-rater agreement: a measure of how consistently two or more qualified labelers assign the same label to the same record, used to test whether the category definitions are precise enough to apply consistently.
  • Adjudication: the defined process for resolving a disagreement between labelers, performed by someone not among the original labelers of that item.
  • Label set version: a frozen, identified snapshot of all labels used for a specific model build, retained so the model can be reproduced.

5. Procedure

5.1 Write precise category definitions

  1. Define every label category in writing, in plain language, with the boundary condition that distinguishes it from its nearest neighbor category stated explicitly.
  2. For every category, include at least <<FILL: number, e.g. two>> worked edge-case examples: real or realistic records that sit close to the boundary, with the correct label and the reasoning stated.
  3. Route the definitions for review by a second SME who did not write them, specifically to find ambiguity before labeling starts, not after.
  4. Version the definitions document; any change to a definition after labeling has started requires re-assessing whether previously applied labels remain valid under the new definition.

5.2 Qualify the labelers

  1. Define the qualification basis for a labeler: role, experience, or a training and assessment record, appropriate to the judgment being made.
  2. Record each labeler’s qualification on file before they label any record that will be used as ground truth.
  3. Where the judgment requires specialized expertise (for example a clinical or analytical judgment), confirm the labeler holds that expertise, not general familiarity with the record type.

5.3 Calibrate before full-scale labeling

  1. Before labeling the full candidate set, have all labelers independently label a small calibration set (<<FILL: number, e.g. 20-30>> records) covering a range of categories including known edge cases.
  2. Discuss disagreements as a group, refine the category definitions if the discussion reveals real ambiguity, and re-version the definitions if changed.
  3. Do not proceed to full-scale labeling until the calibration round shows labelers converging on a shared understanding of the categories.

5.4 Double-label a subset and measure agreement

  1. Select a random subset of the full candidate set for independent double-labeling: at least <<FILL: percentage or count, e.g. 10% or 200 records, whichever is greater>>.
  2. Have two or more qualified labelers label this subset independently, with no visibility into each other’s labels until both are complete.
  3. Calculate inter-rater agreement (see section 6 for the acceptance threshold and section 8 for a worked calculation).
  4. If agreement falls below the pre-set threshold, treat this as a signal the category definitions are ambiguous, not a labeler competence problem in isolation. Return to section 5.1, sharpen the definitions with the specific disagreements as new worked examples, and re-run calibration and double-labeling before proceeding.

5.5 Adjudicate disagreements

  1. For every record where labelers disagreed, route it to the adjudicator, who was not one of the original labelers of that item.
  2. The adjudicator applies the current category definitions, documents the reasoning, and assigns the final label.
  3. Record every adjudication decision, including the original conflicting labels, so the pattern of disagreement remains visible for future definition refinement.

5.6 Version and release the label set

  1. Once all records are labeled, disagreements are adjudicated, and agreement meets the pre-set threshold on the double-labeled subset, freeze the label set with a version identifier.
  2. Record the category definitions version, the labelers involved, the agreement statistic, and the adjudication log as part of the frozen label set’s provenance.
  3. Any subsequent re-labeling, definition change, or addition of new records is a new label set version, not a silent edit to the existing one.

6. Acceptance criteria

  • Every category has a written definition with at least the specified number of worked edge-case examples.
  • Every labeler contributing to the ground truth has a documented qualification basis on file.
  • A calibration round was completed before full-scale labeling, with definitions refined if ambiguity surfaced.
  • At least the specified proportion of the candidate set was double-labeled independently.
  • Inter-rater agreement on the double-labeled subset meets or exceeds <<FILL: pre-set threshold, e.g. 85%>>, appropriate to the difficulty and consequence of the judgment; a lower threshold requires a documented, QA-approved rationale.
  • Every disagreement was adjudicated by someone other than the original labelers, with the reasoning recorded.
  • The label set is versioned and its provenance (definitions version, labelers, agreement statistic, adjudication log) is retained.

7. References

FDA, Data Integrity and Compliance With Drug CGMP, Questions and Answers (Guidance for Industry, December 2018), for the Attributable and Accurate expectations a label carries as a GxP data point. 21 CFR Part 11 and EU GMP Annex 11, for electronic-record controls over the label set and its provenance. MHRA GxP Data Integrity Guidance and Definitions; PIC/S PI 041 (reference by title; describe, do not paste). ICH Q9, Quality Risk Management, for sizing labeler qualification and agreement threshold to the consequence of the labeled judgment.

Confirm the current version and clause numbers of each reference before issue.

8. Worked inter-rater agreement calculation

Percent agreement is the simplest and most transparent measure: the count of records where both labelers assigned the same label, divided by the total number of double-labeled records.

Example: 250 records were double-labeled by two qualified reviewers. The reviewers assigned the same label on 215 of the 250 records.

Percent agreement = 215 / 250 = 0.86, or 86 percent.

Against a pre-set threshold of 85 percent, this result passes. Where the categories are more than two, or where chance agreement is a real concern (a small number of categories with an uneven base rate), consider a chance-corrected statistic such as Cohen’s kappa alongside percent agreement, and set its acceptance threshold separately, since kappa is not on the same 0-100 percent scale and a kappa of 0.6 is not comparable to 60 percent agreement.

9. Record generated: labeling and inter-rater agreement record

FieldEntry
Labeling program / model<<FILL>>
Category definitions version<<FILL>>
Labelers and qualification basis<<FILL>>
Calibration round outcome<<FILL>>
Double-labeled subset size<<FILL>>
Agreement statistic and value<<FILL>>
Threshold metYes / No
Adjudication log reference<<FILL>>
Frozen label set version<<FILL>>
Prepared by (name, signature, date)<<FILL>>
QA approval (name, signature, date)<<FILL>>

10. Revision history

VersionDateAuthorSummary of change
<<FILL: 1.0>><<FILL: date>><<FILL: author>>Initial issue.

11. Approvals

RoleNameSignatureDate
Author<<FILL>>
Reviewer (Data Science)<<FILL>>
Approver (Quality Assurance)<<FILL>>

Filled specimen

The following shows the labeling and inter-rater agreement record completed for an illustrative deviation-criticality labeling program feeding a triage-assist model. Numbers are illustrative; replace them with your own.

FieldEntry
Labeling program / modelDeviation Criticality Triage Assist, ground truth v2
Category definitions versionDEF-DEV-CRIT-v2, revised after v1 calibration showed “Major” and “Critical” boundary ambiguity
Labelers and qualification basisA. Chen (QA, 6 years deviation review), T. Marsh (QA, 4 years, completed deviation-classification competency assessment 2026-05)
Calibration round outcome25-record calibration round; 4 disagreements, all on the Major/Critical boundary; definitions revised to v2 with 3 new worked edge-case examples before full labeling
Double-labeled subset size250 of 1,020 candidate records (24.5%)
Agreement statistic and valuePercent agreement, 86% (215/250); Cohen’s kappa 0.79
Threshold metYes (threshold 85% percent agreement)
Adjudication log referenceADJ-DEV-CRIT-2026-08, 35 records adjudicated by S. Whitfield (QA, senior reviewer, not among original labelers)
Frozen label set versionDEV-CRIT-LABELS-2026-08-14-v2
Prepared byA. Chen, 14 August 2026
QA approvalS. Whitfield, 15 August 2026

The 86 percent agreement, calculated exactly as shown in section 8 (215 correct matches out of 250 double-labeled records), is what let the team proceed with confidence that the categories, after their v2 revision, were precise enough for a model to learn a consistent rule from. The v1 attempt had failed calibration at 71 percent, which is why the definitions were rewritten before any full-scale labeling began.

Common inspection findings this SOP prevents

  • Labels used to train a model with no written category definitions, so no one can say what “Critical” actually meant to the person who assigned it.
  • Labelers with no documented qualification basis for the judgment they were making.
  • No double-labeling and no inter-rater agreement measurement, so label consistency is simply assumed rather than demonstrated.
  • Disagreements resolved informally, with no adjudication record, so the same ambiguity resurfaces in every labeling round.
  • A label set used to train a model with no version identifier, so the exact labels the model learned from cannot be reconstructed later.

How to adapt this SOP

  1. Set your document number, owner, and effective date in the header.
  2. Set the double-labeling proportion in section 5.4 and the agreement threshold in section 6 to match the consequence of the judgment; a higher-consequence label (for example, one that could trigger a regulatory reportability decision) warrants a larger double-labeled sample and a higher threshold than a low-consequence one.
  3. Decide whether percent agreement alone is sufficient or whether a chance-corrected statistic like Cohen’s kappa is warranted, based on your number of categories and their base rates, and set both thresholds explicitly if you use both.
  4. Point the cross-references in section 2 to your source-record procedures for any labels drawn from existing structured categorization rather than fresh SME labeling.
  5. Confirm every regulation in section 7 against the current published version before issue.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.