Independent and not affiliated with the FDA, MHRA, ISPE, PDA, or any agency. Get the appgoutham@madhadi.com
madhadi.comData Integrity & GxP Quality
Browse all topics → Articles Templates & Procedures Learning paths GlossaryScenariosToolsRegulatory ReferencesLearning PathsTopics About Start here
Specification Plug-and-play starting point AI & Automation

Specification: LLM Golden Dataset (Evaluation Reference Set)

A plug-and-play specification for the golden dataset behind a GxP LLM qualification: numbered requirements for coverage, per-item schema, labeler qualifications and inter-rater agreement, statistical sizing, development and locked-evaluation splits, versioning and hashing, and refresh, with a filled specimen.

Document type: Specification

Read and copy the template below into your own quality system. It is a generic starting point for your own internal use, provided as is, with no warranty; see the Terms and License. Adopting it does not by itself create compliance.

This is a ready-to-use specification for the golden dataset, the reference set of inputs and known-good expected behavior that is the test evidence for a generative system. An inspector will ask how it was built; these numbered requirements are what a defensible answer looks like. Replace each <<FILL: ...>> placeholder with your own specifics. A filled specimen follows. Verify each cited regulation against the current source before you rely on it.

Document control header

FieldEntry
Specification titleLLM Golden Dataset for <<FILL: use case>>
Document number<<FILL: SPEC-ID>>
Version<<FILL: version>>
System / use case<<FILL: name>>
Owner<<FILL: Data Steward / role>>
Applies to<<FILL: the qualification this dataset supports>>

1. Purpose and scope

This specification defines the requirements for the golden dataset used to qualify and monitor <<FILL: system>>. It governs what the dataset must contain, how items are labeled, how it is sized, how it is split, and how it is versioned and refreshed, so the metrics computed on it are credible and reproducible.

2. Requirements

Requirements are numbered and testable. Each is verified during dataset review before the qualification run.

2.1 Representativeness

  • GD-01 Inputs shall be drawn from, or built to mirror, the actual production population: real input texts, real questions, and real document layouts, not synthetic-only content that lacks the mess of real data.
  • GD-02 The dataset shall cover the common cases and deliberately include hard and rare ones: ambiguous inputs, sources that do not contain the answer, conflicting sources, very long inputs, unusual formatting, and out-of-scope prompts.
  • GD-03 The dataset shall include unanswerable items where the correct behavior is abstention, and adversarial items including instructions injected into a source document.

2.2 Per-item schema

  • GD-04 Each item shall carry: a unique item ID; the input (prompt and any source documents); the expected behavior (a reference answer, a set of required facts for a summary, or the expectation of abstention); the supporting source spans for groundedness scoring; and a category tag (answerable, unanswerable, adversarial, edge).
  • GD-05 For summary or open-ended items, the expected behavior shall define the required facts that must appear rather than a single golden text, since many outputs are acceptable.

2.3 Labeling and qualification

  • GD-06 Reference answers and source spans shall be written or approved by qualified subject matter experts under a documented labeling SOP, with labeler qualifications recorded, exactly as for any GxP determination.
  • GD-07 Where judgment is involved, at least two reviewers shall label a defined subset; inter-rater agreement shall be measured and recorded, and disagreements resolved by a defined method (for example a third reviewer).

2.4 Sizing

  • GD-08 The dataset shall be sized so the acceptance thresholds are observable and bounded. For a near-zero failure tolerance, the evaluation split shall be large enough that the metric can be estimated with a stated confidence interval; a set too small to observe the claimed failure rate is not acceptable.
  • GD-09 The sizing rationale shall be recorded, including the target threshold, the resulting minimum size, and any deliberate stress-weighting of hard cases (declared as such rather than presented as representative of frequency).

2.5 Splits and leakage control

  • GD-10 The dataset shall be split into a development split (used to tune prompts and retrieval) and a locked evaluation split (touched once to produce the qualification number).
  • GD-11 The locked evaluation split shall never be shown to the people tuning prompts and shall never be used to select the model; tuning against it is the generative analogue of test-set leakage and inflates the reported score.

2.6 Version control and integrity

  • GD-12 The dataset shall be under version control with a recorded hash, so a qualification result references the exact dataset version it was produced under.
  • GD-13 Any change to the dataset shall be a controlled change with its own record; a changed dataset is a changed test basis.

2.7 Refresh

  • GD-14 The dataset shall be reviewed and extended on a defined cadence so it keeps reflecting the live input population; a confirmed regression case (for example a real hallucination found in production) shall be added as a permanent item.

3. Acceptance criteria (dataset review)

  • Every requirement in section 2 is met and evidenced.
  • Coverage includes answerable, unanswerable, adversarial, and edge categories.
  • Labeler qualifications and inter-rater agreement are on file.
  • The sizing rationale ties the size to the thresholds with a confidence interval.
  • Development and locked splits are separated, hashed, and versioned.

4. References

21 CFR Part 11 (electronic records) and ALCOA+ expectations for the records the system produces. ICH Q9(R1) (Quality Risk Management) for sizing evaluation effort to risk. GAMP 5 (second edition) risk-based validation principles, applied to the AI-specific evidence shift; described in original wording, not reproduced. FDA guidance on Computer Software Assurance for Production and Quality Management System Software (current version), for the risk-based assurance approach.

Confirm the current version of each reference before you rely on it.

5. Filled specimen

Illustrative; replace with your own. Golden dataset for a retrieval-QA assistant answering specification questions from a controlled corpus.

AttributeEntry
Total items400
Answerable (with source spans)280
Unanswerable (expect abstention)80
Adversarial (out-of-scope or injection)40
LabelersTwo qualified SMEs; inter-rater agreement 0.91 on a 60-item overlap; disagreements resolved by a third reviewer
Sizing rationaleHallucination tolerance 0.5 percent; evaluation split of 280 supports a bounded estimate with a reported 95 percent interval; hard cases stress-weighted and declared
Splits120 development, 280 locked evaluation, hashed and version controlled
RefreshQuarterly review; the one production hallucination found post-release added as a permanent regression item

This set does the three things a weak set does not: it includes abstention and adversarial cases (not just easy answerable ones), it records who labeled it and how well they agreed, and it keeps the evaluation split locked away from prompt tuning. Those are the properties that make the metrics computed on it mean something in production.

Common inspection findings this specification prevents

  • A tiny, synthetic, all-easy test set with no abstention cases, used both to tune prompts and to report the score.
  • A reported groundedness or accuracy number with no stated basis, resting on too few items to support the claim.
  • Prompt tuning against the same items used to report the qualification metric.
  • A golden dataset with no version, no hash, and no labeler qualifications, so the test evidence cannot be reproduced or trusted.

How to adapt this specification

  1. Set your document number, use case, and owner in the header.
  2. Set the coverage counts and the sizing to your own thresholds and risk; record the rationale.
  3. Reference your real labeling SOP and record labeler qualifications and inter-rater agreement.
  4. Put the dataset and its hash under the same version control as the rest of the system configuration.
  5. Confirm every regulation in the references against the current published version before you rely on it.
Use madhadi.com as an app Full screen, works offline, one tap from your home screen.