This is a ready-to-use SOP. Replace every <<FILL: ...>> placeholder with your own specifics, set your document numbers and dates, and route it through your normal document control, review, and approval. A worked filled specimen, including an inter-rater agreement calculation, follows the template. Verify each cited regulation against the current source before you rely on it.
Document control header
| Field | Entry |
|---|---|
| Document title | AI/ML Ground Truth Labeling |
| Document number | <<FILL: SOP-ID, e.g. SOP-AI-014>> |
| Version | <<FILL: version, e.g. 1.0>> |
| Effective date | <<FILL: effective date>> |
| Supersedes | <<FILL: prior version or "New">> |
| Document owner | <<FILL: role, e.g. Quality Assurance / Data Science Lead>> |
| Applies to | <<FILL: sites / departments / model types in scope>> |
1. Purpose
This procedure defines how <<FILL: COMPANY NAME>> establishes ground truth, the labeled examples an AI or machine learning model learns from, so that labels used to train a GxP-relevant model reflect a defensible, consistent, qualified judgment rather than one analyst’s unrecorded habit. A label is a data point the same way a measurement is, and it carries the same expectation of being attributable, accurate, and reproducible.
2. Scope
This procedure applies whenever a person assigns a label, category, or judgment-based outcome to a record for the purpose of training, validating, or testing an AI/ML model used in a GxP-relevant workflow at the sites in the header. It covers new labeling programs and the re-use of existing structured categorization (for example a historical QMS category) as ground truth. It does not apply where the label is an objective downstream outcome (a measured result, a batch’s actual disposition) rather than a human judgment; those are governed by the source record’s own procedure.
3. Responsibilities
| Role | Responsibility |
|---|---|
| Labeling program owner | Writes and maintains the category definitions, selects and qualifies labelers, and owns the labeling SOP’s fitness for purpose. |
| Qualified labelers (SME) | Apply labels per the definitions in section 5.1; participate in double-labeling and calibration sessions. |
| Adjudicator | A senior reviewer, not among the original labelers of the disputed item, who resolves disagreements per section 5.5. |
| Data Science | Consumes the versioned label set; reports back on any pattern of model error that suggests a label-definition problem. |
| Quality Assurance | Reviews and approves the labeling SOP, the labeler qualification records, and the inter-rater agreement result before the label set is used to train a model. |
4. Definitions
- Ground truth: the set of labels the organization trusts enough to teach a model from, established under this procedure.
- Inter-rater agreement: a measure of how consistently two or more qualified labelers assign the same label to the same record, used to test whether the category definitions are precise enough to apply consistently.
- Adjudication: the defined process for resolving a disagreement between labelers, performed by someone not among the original labelers of that item.
- Label set version: a frozen, identified snapshot of all labels used for a specific model build, retained so the model can be reproduced.
5. Procedure
5.1 Write precise category definitions
- Define every label category in writing, in plain language, with the boundary condition that distinguishes it from its nearest neighbor category stated explicitly.
- For every category, include at least
<<FILL: number, e.g. two>>worked edge-case examples: real or realistic records that sit close to the boundary, with the correct label and the reasoning stated. - Route the definitions for review by a second SME who did not write them, specifically to find ambiguity before labeling starts, not after.
- Version the definitions document; any change to a definition after labeling has started requires re-assessing whether previously applied labels remain valid under the new definition.
5.2 Qualify the labelers
- Define the qualification basis for a labeler: role, experience, or a training and assessment record, appropriate to the judgment being made.
- Record each labeler’s qualification on file before they label any record that will be used as ground truth.
- Where the judgment requires specialized expertise (for example a clinical or analytical judgment), confirm the labeler holds that expertise, not general familiarity with the record type.
5.3 Calibrate before full-scale labeling
- Before labeling the full candidate set, have all labelers independently label a small calibration set (
<<FILL: number, e.g. 20-30>>records) covering a range of categories including known edge cases. - Discuss disagreements as a group, refine the category definitions if the discussion reveals real ambiguity, and re-version the definitions if changed.
- Do not proceed to full-scale labeling until the calibration round shows labelers converging on a shared understanding of the categories.
5.4 Double-label a subset and measure agreement
- Select a random subset of the full candidate set for independent double-labeling: at least
<<FILL: percentage or count, e.g. 10% or 200 records, whichever is greater>>. - Have two or more qualified labelers label this subset independently, with no visibility into each other’s labels until both are complete.
- Calculate inter-rater agreement (see section 6 for the acceptance threshold and section 8 for a worked calculation).
- If agreement falls below the pre-set threshold, treat this as a signal the category definitions are ambiguous, not a labeler competence problem in isolation. Return to section 5.1, sharpen the definitions with the specific disagreements as new worked examples, and re-run calibration and double-labeling before proceeding.
5.5 Adjudicate disagreements
- For every record where labelers disagreed, route it to the adjudicator, who was not one of the original labelers of that item.
- The adjudicator applies the current category definitions, documents the reasoning, and assigns the final label.
- Record every adjudication decision, including the original conflicting labels, so the pattern of disagreement remains visible for future definition refinement.
5.6 Version and release the label set
- Once all records are labeled, disagreements are adjudicated, and agreement meets the pre-set threshold on the double-labeled subset, freeze the label set with a version identifier.
- Record the category definitions version, the labelers involved, the agreement statistic, and the adjudication log as part of the frozen label set’s provenance.
- Any subsequent re-labeling, definition change, or addition of new records is a new label set version, not a silent edit to the existing one.
6. Acceptance criteria
- Every category has a written definition with at least the specified number of worked edge-case examples.
- Every labeler contributing to the ground truth has a documented qualification basis on file.
- A calibration round was completed before full-scale labeling, with definitions refined if ambiguity surfaced.
- At least the specified proportion of the candidate set was double-labeled independently.
- Inter-rater agreement on the double-labeled subset meets or exceeds
<<FILL: pre-set threshold, e.g. 85%>>, appropriate to the difficulty and consequence of the judgment; a lower threshold requires a documented, QA-approved rationale. - Every disagreement was adjudicated by someone other than the original labelers, with the reasoning recorded.
- The label set is versioned and its provenance (definitions version, labelers, agreement statistic, adjudication log) is retained.
7. References
FDA, Data Integrity and Compliance With Drug CGMP, Questions and Answers (Guidance for Industry, December 2018), for the Attributable and Accurate expectations a label carries as a GxP data point. 21 CFR Part 11 and EU GMP Annex 11, for electronic-record controls over the label set and its provenance. MHRA GxP Data Integrity Guidance and Definitions; PIC/S PI 041 (reference by title; describe, do not paste). ICH Q9, Quality Risk Management, for sizing labeler qualification and agreement threshold to the consequence of the labeled judgment.
Confirm the current version and clause numbers of each reference before issue.
8. Worked inter-rater agreement calculation
Percent agreement is the simplest and most transparent measure: the count of records where both labelers assigned the same label, divided by the total number of double-labeled records.
Example: 250 records were double-labeled by two qualified reviewers. The reviewers assigned the same label on 215 of the 250 records.
Percent agreement = 215 / 250 = 0.86, or 86 percent.
Against a pre-set threshold of 85 percent, this result passes. Where the categories are more than two, or where chance agreement is a real concern (a small number of categories with an uneven base rate), consider a chance-corrected statistic such as Cohen’s kappa alongside percent agreement, and set its acceptance threshold separately, since kappa is not on the same 0-100 percent scale and a kappa of 0.6 is not comparable to 60 percent agreement.
9. Record generated: labeling and inter-rater agreement record
| Field | Entry |
|---|---|
| Labeling program / model | <<FILL>> |
| Category definitions version | <<FILL>> |
| Labelers and qualification basis | <<FILL>> |
| Calibration round outcome | <<FILL>> |
| Double-labeled subset size | <<FILL>> |
| Agreement statistic and value | <<FILL>> |
| Threshold met | Yes / No |
| Adjudication log reference | <<FILL>> |
| Frozen label set version | <<FILL>> |
| Prepared by (name, signature, date) | <<FILL>> |
| QA approval (name, signature, date) | <<FILL>> |
10. Revision history
| Version | Date | Author | Summary of change |
|---|---|---|---|
<<FILL: 1.0>> | <<FILL: date>> | <<FILL: author>> | Initial issue. |
11. Approvals
| Role | Name | Signature | Date |
|---|---|---|---|
| Author | <<FILL>> | ||
| Reviewer (Data Science) | <<FILL>> | ||
| Approver (Quality Assurance) | <<FILL>> |
Filled specimen
The following shows the labeling and inter-rater agreement record completed for an illustrative deviation-criticality labeling program feeding a triage-assist model. Numbers are illustrative; replace them with your own.
| Field | Entry |
|---|---|
| Labeling program / model | Deviation Criticality Triage Assist, ground truth v2 |
| Category definitions version | DEF-DEV-CRIT-v2, revised after v1 calibration showed “Major” and “Critical” boundary ambiguity |
| Labelers and qualification basis | A. Chen (QA, 6 years deviation review), T. Marsh (QA, 4 years, completed deviation-classification competency assessment 2026-05) |
| Calibration round outcome | 25-record calibration round; 4 disagreements, all on the Major/Critical boundary; definitions revised to v2 with 3 new worked edge-case examples before full labeling |
| Double-labeled subset size | 250 of 1,020 candidate records (24.5%) |
| Agreement statistic and value | Percent agreement, 86% (215/250); Cohen’s kappa 0.79 |
| Threshold met | Yes (threshold 85% percent agreement) |
| Adjudication log reference | ADJ-DEV-CRIT-2026-08, 35 records adjudicated by S. Whitfield (QA, senior reviewer, not among original labelers) |
| Frozen label set version | DEV-CRIT-LABELS-2026-08-14-v2 |
| Prepared by | A. Chen, 14 August 2026 |
| QA approval | S. Whitfield, 15 August 2026 |
The 86 percent agreement, calculated exactly as shown in section 8 (215 correct matches out of 250 double-labeled records), is what let the team proceed with confidence that the categories, after their v2 revision, were precise enough for a model to learn a consistent rule from. The v1 attempt had failed calibration at 71 percent, which is why the definitions were rewritten before any full-scale labeling began.
Common inspection findings this SOP prevents
- Labels used to train a model with no written category definitions, so no one can say what “Critical” actually meant to the person who assigned it.
- Labelers with no documented qualification basis for the judgment they were making.
- No double-labeling and no inter-rater agreement measurement, so label consistency is simply assumed rather than demonstrated.
- Disagreements resolved informally, with no adjudication record, so the same ambiguity resurfaces in every labeling round.
- A label set used to train a model with no version identifier, so the exact labels the model learned from cannot be reconstructed later.
How to adapt this SOP
- Set your document number, owner, and effective date in the header.
- Set the double-labeling proportion in section 5.4 and the agreement threshold in section 6 to match the consequence of the judgment; a higher-consequence label (for example, one that could trigger a regulatory reportability decision) warrants a larger double-labeled sample and a higher threshold than a low-consequence one.
- Decide whether percent agreement alone is sufficient or whether a chance-corrected statistic like Cohen’s kappa is warranted, based on your number of categories and their base rates, and set both thresholds explicitly if you use both.
- Point the cross-references in section 2 to your source-record procedures for any labels drawn from existing structured categorization rather than fresh SME labeling.
- Confirm every regulation in section 7 against the current published version before issue.