Most data integrity programs focus on what happens at the point of data generation: the analyst runs a test, the instrument records a result, the operator signs the batch record step. That is the visible part. The failure modes that actually drive warning letters are often downstream, in how data is processed, reviewed, transferred, stored, and eventually retrieved.
This article covers the full data lifecycle and the specific integrity risks at each stage. It applies across the GxPs: a chromatography result in a QC lab, a sensor trace from a manufacturing historian, a case report form value in a clinical trial, and a complaint record in a device quality system all move through the same arc. It is written for practitioners who already understand ALCOA+ and are thinking about how to build governance around the entire life of a data asset, not just its creation. If you are still building the foundation, start with data integrity foundations and ALCOA+ in detail, then come back here.
The Lifecycle Stages
Why the lifecycle framing exists at all
Regulators moved to a lifecycle model because integrity findings kept appearing in places that a creation-only control set did not cover: deleted raw files, reprocessed chromatograms, untested backups, records that could not be produced during an inspection. The FDA December 2018 data integrity guidance, titled “Data Integrity and Compliance With Drug CGMP: Questions and Answers,” treats integrity as a property across the full data life cycle, from creation through modification, processing, maintenance, archival, retrieval, transmission, and disposition. MHRA guidance lands on the same scope from the regulator’s side of the desk: nothing about a data set’s evidentiary weight resets partway through its life, so a firm cannot treat the obligation as satisfied once a result is generated and signed off. The expectation reaches every later thing that happens to that data, being read, transformed, filed away, dug back out years afterward, and eventually cleared out under a disposal schedule, and MHRA is explicit that none of those later moments falls outside the same duty that applied at the moment of capture. Practically, that closes a loophole a lot of programs lean on without meaning to: treating the lifecycle idea as a checklist for the lab bench while quietly assuming that archives and disposal are a records-management afterthought governed by looser rules.
The MHRA’s “GXP Data Integrity Guidance and Definitions” (March 2018) carries that lifecycle definition, and PIC/S guidance PI 041-1 (July 2021), “Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments,” extends it with detailed expectations for each stage. The WHO has issued a parallel annex on good data and record management practices, and in the clinical world ICH E6(R3) Good Clinical Practice (Step 4, January 2025, which superseded E6(R2)) carries the same principle for trial data. The documents agree on the core point: integrity is a property of the whole lifecycle, not a checkbox at creation.
A compendial revision is moving the same direction. USP General Chapter <1029>, currently in force under the title “Good Documentation Guidelines” (a 2018 text focused on documentation mechanics), was proposed for a substantial rewrite and retitling to “Good Documentation Guidelines and Data Integrity,” published in the Pharmacopeial Forum for public comment with the comment period closing 30 September 2025. The proposal folds ALCOA+ principles and a data governance framing directly into the chapter and flags that a separate, dedicated data governance general chapter is planned as a future addition. As of this writing the revision has not been adopted into the official USP-NF text, so the 2018 documentation-only chapter remains current; treat the retitled version as a direction of travel worth watching, not yet a citable standard. General chapters numbered <1000> and above are informational rather than mandatory by default, so <1029> supports a documentation and data governance program rather than imposing an independent compendial requirement, but it is frequently the clearest single reference GMP auditors and trainers point to for what “good documentation” means in practice.
The legal anchor in the US predates all of the guidance. 21 CFR 211.180 and 211.194 require complete records and that they be retained and available; 21 CFR Part 11 governs electronic records and signatures; the EU equivalents sit in EudraLex Volume 4 Chapter 4 (Documentation) and Annex 11 (Computerised Systems). The guidance documents interpret those rules across the lifecycle, they do not invent new authority.
The seven stages
In practice the lifecycle maps to seven stages for most GxP data:
- Generation: data is created by an instrument, a human, or a system.
- Processing: raw data is transformed (integration, calculation, aggregation).
- Review: data is assessed against specifications and approved.
- Reporting: data is summarized and included in decisions or submissions.
- Retention: data is stored and maintained.
- Retrieval: data is accessed for review, investigation, or inspection.
- Archival and disposition: data is transferred to long-term storage or destroyed per schedule.
Each stage has distinct integrity risks. A program that only addresses generation will have gaps at processing, review, and archival that inspectors will find. The table below summarizes the dominant risk and the primary control at each stage.
| Stage | Dominant integrity risk | Primary control |
|---|---|---|
| Generation | Wrong thing treated as the original; uncontrolled paper | Define the raw data; restrict local saves |
| Processing | Reintegration without justification; transcription error | Locked methods; reason codes; validated calculations |
| Review | Reviewer never sees the level where manipulation hides | Audit trail review built into the release flow |
| Reporting | Transcription and transfer corruption; selective reporting | Verified transfers; completeness rule |
| Retention | Legibility loss; untested backups; format obsolescence | Restoration testing; format migration plan |
| Retrieval | Records exist but cannot be produced promptly | Searchable archive; tested, timed retrieval |
| Disposition | Premature destruction; destruction under legal hold | Approved retention schedule; legal hold process |
A useful way to govern this is to require, for every GxP system, that someone can name the control at each of the seven rows. If a row is blank, that is the next finding waiting to happen.
Stage 1: Generation
Generation is where the primary record is created. The key question at this stage is: what is the original?
For paper-based systems, the original is the first recorded version: the raw entry in the notebook, the printed chart recorder trace, the manual logbook line. A photocopy is not the original. Good documentation behavior governs this point, and it is worth reading good documentation practices alongside this one.
For electronic systems, FDA’s position is explicit: the electronic raw data is the original, not any printout derived from it. A chromatogram printed from a chromatography data system is a representation of the original. The raw data file, the binary or structured file containing the acquisition data and audit trail, is the original that must be retained. This has been the basis of many warning letters where firms retained printed chromatograms but deleted the underlying electronic files.
Static vs dynamic records
The distinction matters for understanding what “original” means, and it is one of the most tested concepts in an inspection. A dedicated treatment lives in static and dynamic records and true copies; the short version follows.
- A static record cannot be meaningfully interacted with after creation: a PDF, a printed report, a scanned image. Its content is fixed at the time of creation.
- A dynamic record can be processed, requeried, and reanalyzed: a chromatography data system raw data file, an electronic laboratory notebook dataset, a LIMS result record, a manufacturing execution system batch object. The original dynamic record preserves all the processing metadata and allows reconstruction of every analytical decision.
For dynamic records, retaining only a static export (a PDF of the result) loses the ability to verify that the result correctly represents the raw data. The MHRA guidance is blunt here: a static printout of a dynamic record is not a true copy, because it does not carry the metadata and reprocessing capability of the source.
True copy and the original/copy trap
A true copy is a copy of original data that has been verified to preserve the full content and meaning, including all metadata and, for dynamic records, the dynamic nature of the source. If a true copy is created and verified, the original may in some defined cases be retained as the copy. The trap is treating an export as a true copy when it has silently dropped audit trail entries, timestamps, or reprocessing capability. Verification, not assumption, is what makes a copy “true,” and the verification itself must be recorded (who verified, against what, when).
How to define the raw data, step by step
The single most useful generation control is a written definition of raw data for each system. Without it, “the original” is an opinion. The procedure:
- List every GxP system that generates or holds data.
- For each, identify the file or record that is created first and contains the complete result plus its metadata and audit trail.
- State explicitly whether that record is static or dynamic.
- Define where it is stored, who owns it, and the retention period.
- State what is not the raw data (for example, the PDF printout, the local working copy) so there is no ambiguity during review or destruction.
- Approve the definition through the quality system and reference it in the system’s validation package and standard operating procedures.
Acceptance criteria. A reviewer or inspector can ask “what is the raw data for this test?” and be shown a single, approved answer that matches what the system actually retains. The retention configuration of the system enforces it (raw files are not deletable by the analyst). The definition covers metadata and audit trail, not just result values.
A worked example
An analyst acquires a liquid chromatography injection. The chromatography data system writes a result set: the chromatogram, the integration events, the method, the sequence, the audit trail, the user and instrument identifiers, and the acquisition timestamp. That bundle is the original dynamic record. If site procedure says “print the result, file the printout, then purge the project from the system to free disk space,” the site has destroyed the original and kept a static representation. Even where nobody set out to mislead, the underlying record can no longer be rebuilt from what is left behind, and that loss of reconstructability is exactly the kind of shortfall an investigator would typically zero in on and could cite as a deficiency against the recordkeeping expectations.
A concrete control at generation: configure the system so analysts cannot save acquisitions to a local drive or a working folder outside the audited project structure. Local “scratch” saves are a recurring root cause because they create a copy that lives outside the audit trail, which is exactly the gap that enables a hidden “trial” injection. The same logic applies outside the lab. A device test bench that writes results to an engineer’s laptop, or a manufacturing sensor that buffers to a local cache before the historian ingests it, both create un-audited intermediate copies.
Stage 2: Processing
Processing is where raw data is transformed into reportable results. In analytical chemistry, this typically means peak integration, baseline correction, calculation against a standard, and evaluation against a specification limit. In manufacturing it can mean smoothing, scaling, or aggregating a sensor trace. In clinical data management it means edit checks, queries, and derivations. The risks rhyme across all three.
The integrity risks at this stage are well documented in enforcement actions:
Integration parameter manipulation. In chromatography data systems, an analyst can reprocess an injection with different integration parameters, wider peak windows, or different baseline assignments, to change the peak area and therefore the result. When this is done without a documented scientific justification behind it, it is widely treated by inspectors and quality units alike as a serious data integrity concern, and that concern holds whether or not the original data file happens to have been retained. Chromatography data system integrity covers the system-specific controls, and the FDA warning letters patterns article shows how often this appears in enforcement.
Selective reprocessing. Not all reprocessing is illegitimate. Integration methods need to be validated and locked before use. When reprocessing occurs after reviewing a result, the audit trail should capture who initiated the reprocessing, what parameters changed, the previous values, and what justification was provided. Systems that log “reprocessed” without capturing the previous parameter values, or that do not require a reason, are a control gap.
Calculation errors. Manual transcription of instrument readings into calculated results, for example reading a UV absorbance and manually calculating concentration, introduces transcription risk. Validated calculation tools, formula locks in spreadsheets, or direct LIMS interfaces reduce this risk. Spreadsheets used for GxP calculations are a system in their own right and need validation; infrastructure qualification and spreadsheet validation addresses how far that has to go.
Reading the audit trail at this stage
The audit trail at processing should capture original values, changed values, the identity of who made each change, the timestamp, and, for controlled changes, the reason. The signals an experienced reviewer looks for:
- Injections or runs aborted or deleted shortly before a “good” run on the same sample.
- Reintegration that moves a result from just outside specification to just inside it.
- A reason of “operator error” or “no peak” applied repeatedly to inconvenient results.
- System clock changes near the time of acquisition (see time stamps and system clock control).
- Reprocessing performed by a different user from the one who acquired the data, with no second-person rationale.
None of these is proof of falsification on its own. Each is a question that the record should be able to answer. For the technical design that makes these signals visible, see audit trail design and review.
A worked processing audit trail entry
A defensible reprocessing event looks like this in the audit trail:
| Field | Value |
|---|---|
| Record | Sample S-2045, injection 3, assay result |
| Action | Reintegration |
| Old value | Area 14,820; result 99.1% |
| New value | Area 15,060; result 100.7% |
| Reason | Manual baseline correction, shoulder peak not resolved by default method per OOS-2045 |
| Performed by | analyst.jdoe |
| Reviewed by | qa.msmith (second person) |
| Timestamp | 2025-09-12 14:07 site local |
What makes it defensible is that the prior value is captured, the reason references an investigation, and a second person reviewed it. Compare that to an entry that reads only “Reintegrated, area 15,060” with no prior value and no reason. The second entry is the finding.
Stage 3: Review
Review is when a second person, typically QA or a supervising analyst, evaluates the data and decides whether it is acceptable. The integrity risks here are subtle but equally important.
Review scope. Reviewers often see summary data or final result reports, not the underlying raw data. A reviewer who sees “Result: 98.5%, PASS” without access to the raw chromatogram, the audit trail, and the sequence log cannot actually verify data integrity. Effective review requires access to the data at the level where manipulation could have occurred. The batch record review process should give reviewers that depth; see batch record review in GMP.
Audit trail review. FDA expects that audit trails are reviewed as part of the batch release process, not just at an annual periodic review. The 2018 guidance states that “the agency recommends routine scheduled audit trail review based on the complexity of the system and its intended use.” In practice this means QA reviewers need access to audit trail records and need to know what anomalies to look for. Reviewers also need the authority to require source data when something looks off. For the operational mechanics, see operationalizing audit trail review.
Review by exception. For high-volume operations, 100% manual audit trail review is impractical. Risk-based review by exception is acceptable, prioritizing high-criticality systems, high-risk time periods (end of shift, weekend, out-of-specification events), and unusual patterns. The key is that the scope and approach are documented and justified, not that review is quietly eliminated. A defensible review-by-exception design ties the sampling rate to a documented risk assessment and still triggers full review on defined exception events.
Who reviews matters. Review independence is part of the control. The person who generated a result should not be the sole reviewer of the audit trail for that result. Segregation of duties at the review stage is one of the controls inspectors probe, and it connects to access design covered in computerized systems access control and cybersecurity.
A second-person review checklist
A practical reviewer checklist for a single QC result, usable as the spine of a procedure:
- Result value matches the raw data and the reported significant figures.
- The method and sequence used are the approved, locked versions.
- Audit trail shows no unexplained deletions, aborts, or reprocessing.
- Any reprocessing carries a prior value and an approved reason.
- No system clock anomalies near acquisition.
- The analyst is qualified and the account is theirs (no shared login).
- Any out-of-specification or out-of-trend result is linked to an open or closed investigation.
- The reviewer is independent of the person who generated the data.
Acceptance criteria. Each item is answerable from the record, the reviewer signs and dates, and the depth of review (summary only vs full audit trail) matches the documented risk-based policy.
Stage 4: Reporting
Reporting is where data moves outside the generating system, into a batch record summary, a regulatory submission, a stability report, or a clinical study report. This stage introduces transfer risk.
Data transfer integrity. When data moves from one system to another, from a chromatography system to a LIMS, from LIMS to an enterprise resource planning system, from there to a submission template, the transferred values must match the originals. This sounds obvious. In practice, manual transcription is still common, and even automated transfers can corrupt, truncate, or misformat data without detection unless verified.
Verification methods include:
- Checksum or hash comparison between source and destination files.
- Row-count and sum reconciliation for numerical datasets.
- Spot-check review of transferred values against source records.
- Validated electronic interfaces with built-in error detection and exception handling.
Each method has a place. A validated interface is the strongest control because it makes verification continuous rather than a one-time event, but it still needs periodic confirmation that the interface has not drifted, for example after a system upgrade on either side. Treat an interface change as a change control event; see change control for validated systems.
Reporting scope. The completeness principle of ALCOA+ applies at reporting: all results relevant to a quality decision must be reported, not just the passing ones. A batch record that summarizes results without disclosing the out-of-specification result that was “invalidated” is a completeness failure, regardless of whether the invalidation was technically justified. The handling of those results has to follow a documented out-of-specification investigation process, and the investigation outcome travels with the data.
Rounding and reporting precision. A quieter reporting risk is inconsistent rounding and significant-figure handling between the analytical method, the calculation tool, and the final report. Define rounding rules in the method, apply them once, and make sure the reported value is the value the specification is evaluated against. Rounding a result down into specification is a classic finding.
Worked transfer verification
A numeric reconciliation makes “verified transfer” concrete. A LIMS exports 12 stability assay results to a report template:
| Check | Source (LIMS) | Destination (report) | Result |
|---|---|---|---|
| Record count | 12 | 12 | Match |
| Sum of assay values | 1,193.4 | 1,193.4 | Match |
| Min / max | 98.7 / 101.2 | 98.7 / 101.2 | Match |
| Spot-check row 7 | 100.4 | 100.4 | Match |
| Out-of-spec flagged | 0 | 0 | Match |
The reconciliation is signed, dated, and filed with the report. If the destination sum had been 1,093.4, the mismatch would have caught a dropped or truncated value before the report informed a decision.
Stage 5: Retention
Retention is where a lot of programs have quiet failures that do not surface until an inspection, or until someone tries to retrieve a record years later.
Retention periods. 21 CFR 211.180(a) requires that batch production and control records be retained “for at least 1 year after the expiration date of the batch” and, for certain over-the-counter drug products without expiration dating, 3 years after distribution. For records related to a New Drug Application or a Biologics License Application, retention obligations are defined in the relevant regulations and the marketing application commitments. EU GMP, in EudraLex Volume 4 Chapter 4, sets retention of at least one year past batch expiry and at least five years after certification, whichever is longer. For clinical trials, ICH E6(R3) Good Clinical Practice (Step 4, January 2025) sets essential record retention tied to marketing application timelines, carrying over the principle from the superseded E6(R2). For combination products that include a device constituent, the device records under 21 CFR 820 (now the Quality Management System Regulation aligning with ISO 13485) are generally retained for the design and expected life of the device but not less than two years from release. Many firms standardize on a single longer period to avoid tracking product-specific dates record by record. Whatever period is chosen, it has to be defined in a procedure and applied consistently.
Retention media and format. Records must remain legible and retrievable throughout the retention period. This creates specific obligations:
- Paper records must be stored under conditions that prevent deterioration. Thermal paper fades and is a known problem for long retention periods, so a verified true copy is often made at the time of generation.
- Electronic records must be backed up, with restoration tested periodically. A backup that has never been successfully restored is not functionally a backup.
- Proprietary file formats require either a long-term license for the rendering software or a migration plan. A raw data file that can only be opened by one software version, and that version is no longer supported, is a legibility problem waiting to surface.
The validation of backup and restore is its own discipline; see backup, restore, and disaster recovery validation.
System decommissioning. When a software system is retired, all GxP records it contains must either be migrated to a replacement system through a validated migration or archived in a format that remains accessible. Firms frequently underestimate the scope and decommission systems without confirming that all records transferred and that the archive is searchable. Migrating result values while losing the audit trail and metadata is the most common way this goes wrong. Data migration validation walks through how to verify a migration so the archived record is still complete and attributable.
A retention schedule fragment
A retention schedule should be specific enough that a records owner can act without interpretation:
| Record type | Regulatory basis | Retention period | Owner | Media |
|---|---|---|---|---|
| Batch production record | 21 CFR 211.180(a) | Expiry + 1 yr, min 10 yr | QA Operations | Validated EDMS |
| Chromatography raw data | 21 CFR 211.194; Part 11 | Expiry + 1 yr, min 10 yr | QC Lab Systems | CDS archive (dynamic) |
| Stability data | 21 CFR 211.166 | Life of program + 1 yr | Stability | LIMS archive |
| Validation records | Company procedure (per validation policy) | System life + retention | Validation | Validated EDMS |
| Complaint records (device) | 21 CFR 820 / QMSR | Device life, min 2 yr | Device QA | QMS |
Acceptance criteria. Every GxP record type maps to a row, each row cites a real regulatory basis, an owner is named, and the storage system enforces the period (records are not deletable before it expires).
Stage 6: Retrieval
A record that exists but cannot be retrieved is not functionally available. FDA expects records to be producible promptly during an inspection, not after a multi-day retrieval effort from an off-site archive. The 2018 guidance ties this to 21 CFR 211.180(c), which requires that records be readily available for authorized inspection.
Archive accessibility. For electronic archives, the archive must be searchable, not just a dump of files. A LIMS archive that requires IT to restore a backup tape before a ten-year-old batch record can be read is a retrieval problem, even if the data technically still exists. Test retrieval on a defined schedule using realistic queries, for example “produce the full result set, including audit trail, for a named batch from a retired system,” and time it. Inspection readiness depends on this working under pressure; see FDA inspection readiness.
Third-party records. Records generated by contract research organizations, contract manufacturers, and clinical sites remain the sponsor’s responsibility. Contractual access rights are necessary but not sufficient. The sponsor needs to have actually exercised those rights and confirmed that the records exist and are retrievable in their original form. This belongs in the quality agreement and in routine oversight; see CDMO oversight and quality agreements.
A retrieval test you can run
A defensible retrieval test, run at least annually per critical archive:
- Pick a batch or study at random, ideally one from a retired system.
- Request the complete record set: results, metadata, and audit trail.
- Time the retrieval from request to a usable, readable record.
- Confirm completeness against the original raw data definition.
- Record the elapsed time and any gaps; remediate gaps through CAPA.
Acceptance criteria. The complete record (including audit trail and metadata) is produced within the time you would need during an inspection (commonly minutes to a few hours, defined in your procedure), it is legible, and nothing is missing relative to the raw data definition.
Stage 7: Archival and Disposition
Disposition, the scheduled destruction of records at the end of their retention period, is often overlooked in data integrity programs. It matters because:
- Disposing of records before their required retention period has fully run cuts directly against the recordkeeping rules and is generally treated as a compliance failure, because the obligation to keep them is still live at the moment they are destroyed.
- Disposing of records relevant to an open regulatory matter (inspection, investigation, litigation) creates serious legal risk.
- Archival media that degrades before the retention period ends must be detected and remediated, not silently lost.
A governance program should include a scheduled review of records nearing end of retention to confirm they are eligible for destruction, a documented destruction record (what was destroyed, when, by whom, under what authorization), and a legal hold process that can suspend disposal when needed. The legal hold has to be able to override the routine schedule, and the override needs to be auditable.
The disposition procedure, step by step
- The system or schedule flags records reaching end of retention.
- The records owner confirms the period has genuinely expired against the approved schedule.
- A legal hold check confirms no open inspection, investigation, recall, or litigation touches the records.
- QA authorizes destruction in writing.
- Destruction is executed and a destruction record is created (record type, quantity, date, method, executor, authorizer).
- The destruction record is itself retained.
Acceptance criteria. No record is destroyed without an authorized destruction record; no record under legal hold is destroyed; the destruction record survives the destroyed data. A common finding is a destruction certificate that lists “various lab records” with no traceability to what was actually destroyed, which is indistinguishable on paper from improper disposal.
Metadata: Part of the Record
One of the most frequently misunderstood aspects of data integrity is the status of metadata.
The FDA 2018 guidance defines metadata as “the contextual information required to understand data,” and gives examples such as a data set that includes the user who acquired it, the date and time of acquisition, the instrument, the link to a particular study or batch, and the processing applied. Timestamps, user IDs, system configuration, instrument identifiers, processing parameters, location, and batch associations are all metadata.
Metadata is not supplementary documentation. It is part of the record. An instrument reading without its acquisition timestamp, instrument ID, and operator attribution is incomplete data. The ALCOA+ requirement for “Complete” applies to metadata as much as to the primary result values, and the requirement for “Attributable” cannot be met without it.
The practical implication: when data is transferred, archived, or exported, metadata must travel with it. An export that strips timestamps and user IDs produces data that can no longer be assessed for integrity. This is a common failure mode in migration projects where the focus is on result values and the metadata is lost in the translation. It is also why a static PDF, which usually carries little or none of the source metadata, fails as a substitute for a dynamic record.
There is a second, subtler point. Metadata can itself be the evidence of a problem. A reanalysis whose only fingerprint is a changed integration parameter buried in the method metadata, a result file with an acquisition time that predates the sample’s logging, a user attribution that points to a shared account: these are detectable only if the metadata is preserved and reviewed. Preserving metadata is what makes the rest of the lifecycle controls work.
What metadata to preserve, by example
A short inventory of metadata categories and why each matters:
| Metadata category | Example fields | Why it matters |
|---|---|---|
| Attribution | User ID, role, signature | Establishes who; required for Attributable |
| Temporal | Acquisition date/time, time zone | Establishes when; sequence and clock integrity |
| Instrument and system | Instrument ID, software version, method ID | Reconstruction and traceability |
| Processing | Integration parameters, calculations, prior values | Detects manipulation; supports reprocessing review |
| Context | Batch/lot, study, sample, project links | Ties data to the decision it supports |
| Audit trail | Change history, reasons, before/after | The core integrity evidence layer |
When you define raw data at generation, define which of these categories the raw record must carry. The test of a “true copy” is whether all of them survive the copy.
Metadata Governance Across Systems
Metadata does not usually stay in the system that generated it. A chromatography result becomes a release value in a LIMS, the LIMS disposition status becomes a batch record entry in a manufacturing execution system or electronic batch record, and the batch context becomes a line item in an ERP quality module. A process historian tag becomes an in-process check in that same batch record. Every one of those hops is a system boundary, and every system boundary is a place where metadata gets narrowed, renamed, or silently reinterpreted.
The failure mode is rarely the loss of the primary result value; systems are built to move that faithfully. What gets lost is the context around it: the instrument identifier, the method version, the timestamp precision, the operator attribution tied to a specific action rather than a generic interface account, and the audit trail entry explaining a change. A receiving system with a narrower metadata schema than the source is a common and often invisible cause of this. An ERP quality module that stores a pass or fail flag and a lot number, for example, may have no field at all for the instrument ID or method version the LIMS held. The data did not fail to transfer; the target system was never built to hold it.
Governing this means naming, for each interface, which system is authoritative for which metadata field, so that a disagreement between two systems has a defined tiebreaker instead of an improvised answer during an inspection. A metadata mapping specification, built and verified per interface the same way a data flow map is built per record type, is the practical control: field by field, source to target, with the transformation stated and the fields that do not survive the hop flagged explicitly rather than discovered later.
Common cross-system handoffs and where metadata typically thins out:
| Handoff | Metadata most at risk | Typical control |
|---|---|---|
| Chromatography or analytical system to LIMS | Instrument ID, method version, integration parameters | The analytical system stays system of record for raw metadata; the LIMS is treated as a summary, not a replacement |
| Process historian to MES/EBR | Tag-to-equipment-to-batch linkage, especially after a tag rename | Tag classification and change control on the historian side, covered in process historian data integrity |
| MES/EBR to ERP | Electronic signature meaning, and the specific step a signature applied to | Interface qualification confirms the signature reference, not only the batch status, carries across |
| Instrument to LIMS by manual entry | Full audit trail context, when an analyst keys a value instead of an interface capturing it | Prefer validated interfaces over manual entry for release-critical data; where manual entry is unavoidable, require second-person verification |
| LIMS or MES to a reporting layer or data warehouse | Everything not needed for the report, by design | Treat the reporting layer as derived data, never the record of record, and keep the link back to source |
None of this is exotic. It is the same completeness and attributability principle from ALCOA+, applied at the seam between systems instead of within one system. MES, EBR, and SCADA data integrity, database integrity and DBA governance, and LIMS implementation and validation each cover one side of these handoffs in depth; this section is about the seam itself. A written metadata governance procedure, naming the authoritative system per field and requiring a documented mapping and verification for every new or changed interface, is what keeps this consistent across an estate instead of solved one integration at a time. See the metadata governance and preservation SOP and the existing interface data mapping verification form for the field-level control.
Building Lifecycle Governance
A mature data lifecycle governance program needs:
- A data flow map for each GxP system: where data enters, how it is processed, where it goes, and who has access at each stage. The map is what turns “we have controls” into “we know where the gaps are.”
- Defined data owners at each stage: who is responsible for integrity at generation, review, retention, and retrieval. Ownership that ends at release is a common blind spot.
- Retention schedules aligned to regulatory requirements, with product-specific and system-specific periods defined and approved.
- Migration and decommissioning procedures that include data integrity verification, complete with metadata, before any system is retired.
- Retrieval testing: periodic, timed confirmation that archived data can actually be retrieved in a reasonable timeframe, including from retired systems.
- Periodic review of the controls themselves, so the program does not silently degrade between inspections; see validation master plan and periodic review.
Roles and responsibilities
Lifecycle governance fails when ownership is fuzzy. A workable split:
| Role | Lifecycle responsibility |
|---|---|
| Data owner (process or lab manager) | Accountable for integrity of their data across all seven stages |
| System owner / IT | Backup, restore, archive, access control, retention enforcement |
| QA | Audit trail review, release decisions, approval of retention and destruction |
| Analysts / operators | Correct generation, no local saves, accurate processing with justified changes |
| Validation | Migration, decommissioning, periodic review of controls |
| Records management | Schedule maintenance, legal holds, destruction records |
The recurring blind spot is the handoff between system owner and data owner. The system owner assumes the data owner reviews audit trails; the data owner assumes IT preserves everything in the archive. Name both responsibilities explicitly. See GxP roles and responsibilities and data governance roles and careers for how these map to a wider quality organization.
Data Flow Mapping as a Formal Deliverable
Everything above assumes someone can actually answer the question “where does this data go, and what happens to it at each stop.” That answer is a data flow map, and it deserves to be treated as a formal, approved deliverable rather than a diagram someone draws for a validation report and never updates again.
A data flow map is data-centric, not system-centric. A single map for, say, a finished-product release result, will typically cross the instrument, the chromatography data system, the LIMS, the batch record system, and the archive, spanning systems that were each validated individually but never looked at together as one path. That combined view is exactly what an inspector is checking for when a data integrity question moves from “is this system validated” to “walk me through what happens to this result.” A map that stops at a system boundary is a map that fails that question.
Building one well means walking it with the people who actually move the data, not reconstructing it from the SOP, because the workaround that never made it into the procedure, the local export, the manual re-key, the shared drive, is exactly the gap the map exists to find. Once built, the map feeds two things directly: data criticality, because a stage where a record can be silently altered or lost deserves more scrutiny than one where it cannot; and audit trail review scope, because the map shows which systems in the path actually need routine review and which are passthrough. Building a GxP data flow map covers the method in full, and the data flow map worksheet with its companion data flow mapping SOP turn the method into a controlled, repeatable deliverable rather than a one-off exercise.
Hybrid Record Risk Across the Lifecycle
The lifecycle framing above assumes a single governing record moving through seven stages. A hybrid record, where the complete account of an activity lives partly on paper and partly in an electronic system, breaks that assumption at every stage, because there are two halves and the framework has to say which one governs and how the two stay in agreement. Hybrid systems and paper-and-electronic records covers hybrids in full; the table below maps the hybrid-specific risk onto the seven lifecycle stages from earlier in this article.
| Stage | Hybrid-specific risk | Control |
|---|---|---|
| Generation | Electronic instrument output and a paper worksheet both exist, and nobody has declared which one governs | Declare the governing record in a hybrid inventory before the system goes into use |
| Processing | A value is transcribed from the electronic side to paper, or the reverse, and drifts silently | Independent verification of every transcribed value, not a glance |
| Review | The reviewer signs the paper and never opens the electronic audit trail behind it | Require the review to start from the electronic side, not the paper |
| Reporting | A paper summary omits an electronic reprocessing or reintegration event the paper author never saw | Cross-reference the report against the electronic audit trail before issue |
| Retention | Paper degrades, thermal print fades, while its electronic counterpart is purged for space; neither half stays complete on its own | Retain both halves for the declared record’s full retention period; certify a true copy of paper before it degrades |
| Retrieval | The paper is filed in one place and the electronic half sits in another system with a different retention clock | A single retrieval index that returns both halves together, not one at a time |
| Disposition | One half of the pair is destroyed on schedule while the other survives, so what remains can no longer be reconciled | Tie the disposition decision to the record pair, not to either half independently |
The pattern across all seven rows is the same: a hybrid needs an explicit answer, at every stage, for which half is authoritative and how the two are kept from silently disagreeing. Programs that only solve this at generation, by declaring a governing record, but not at retention and disposition, end up with a defensible start and an indefensible finish. The hybrid record control and reconciliation SOP and its companion reconciliation checklist build the generation-through-review controls; extend the same discipline to the later stages using the retention and disposition practices in this article.
Legacy System Decommissioning and Migration Integrity
The retention section above flags system decommissioning in a few sentences because it belongs there. It deserves a longer treatment on its own, because decommissioning is where lifecycle programs most often lose data permanently rather than merely risk it. A control gap at generation or review is usually recoverable; a control gap at decommissioning often is not, because the system that could have answered the question is gone.
The recurring root cause is treating decommissioning as an IT project rather than a data integrity event. “Turn off the server, keep a database dump somewhere” satisfies nobody’s actual retention obligation and is exactly the pattern behind findings where a firm cannot produce a record from a system retired years earlier. System decommissioning, data archival, and lawful retention covers the full discipline; the points below are the ones a lifecycle program has to get right regardless of which article or template you start from.
Migrate versus archive. The decision hinges on whether the record set is still dynamic in a way that matters, meaning it may need to be reprocessed, requeried, or reanalyzed, or whether it is genuinely reference-only. Dynamic data flattened into a static export loses exactly the reprocessing capability that made it dynamic in the first place, and that loss is often irreversible once the source system is gone.
Sequence matters more than most teams expect. Verify migration or archival completeness first. Run a retrieval test second, opening real records and confirming they are complete and readable. Only then remove access, and only after that, sanitize or destroy media. Reversing that order, removing access or destroying media before verification is complete, is a common way an otherwise well-planned decommissioning still loses data, because there is no way back once the source is gone.
Migration is itself a validated activity, not a side effect of buying new software. Reconciliation by record count and checksum, plus a full review of the highest-criticality records rather than a thin sample, is what separates a verified migration from an assumed one. Data migration validation and the data migration validation protocol cover the mechanics.
A legacy system that was never properly validated complicates all of this, because you cannot fully trust that what you are migrating or archiving is complete if the system that produced it was never confirmed to work as intended. That gap needs its own assessment, and sometimes its own retrospective validation, before or alongside the decommissioning effort. See retroactive validation and legacy systems.
For the plan itself, the system decommissioning and retirement plan walks the inventory-to-verification-to-access-removal sequence end to end, and the records retention and disposition schedule is what the inventory step should be pulling its retention periods from, rather than each decommissioning project inventing its own. Pair the plan with the records retention, archival, and retrieval verification SOP so the archive this decommissioning creates does not become the next decade’s unread box, and with the archive retrieval and readability test checklist to actually run a retrieval test rather than assume one would work.
Worked Example: One Result Across the Full Lifecycle
The stage-by-stage examples earlier in this article each show one moment in the lifecycle. It is worth also seeing one record move through all seven stages end to end, because that is closer to what an inspector actually asks for: not “show me your generation controls” but “show me what happened to this result, completely, from the day it was created.”
Follow assay result R-5518, an HPLC release test for finished-product batch B-4402.
| Stage | What happens to R-5518 | Integrity control applied |
|---|---|---|
| Generation | The chromatography data system acquires the injection and writes the dynamic raw data file: chromatogram, integration events, method, sequence, audit trail, instrument ID, and acquisition timestamp | The raw data definition names this file, not any printout, as the original; local saves outside the audited project are disabled |
| Processing | The analyst applies the locked, validated integration method; no reintegration is needed | Locked methods and formula-driven calculation remove manual transcription risk |
| Review | QA opens the CDS result and its audit trail, not only the LIMS summary, confirms no unexplained reprocessing, and signs as second-person reviewer | Audit trail review is built into the release step, not deferred to an annual review |
| Reporting | The LIMS exports the result into the batch disposition package and the certificate of analysis; a reconciliation check confirms the reported value matches the CDS source | Verified transfer, count, value, and spot-check reconciliation, before the value is used in a decision |
| Retention | R-5518 is retained in the CDS under the site’s standard retention period, batch expiry plus one year at minimum, with its full metadata and audit trail | Retention period drawn from the approved schedule and enforced by the system, not by an analyst’s discretion |
| Retrieval (year 6) | A product complaint on a later batch triggers a look-back. By this point the original CDS has been decommissioned and replaced; R-5518 now lives in a verified archive built at decommissioning | The decommissioning plan’s migration verification and periodic archive-retrieval test are what make this retrieval possible at all; the archive is opened, and the record and its audit trail are produced within the timeframe the retrieval procedure requires |
| Disposition (year 11) | The retention period expires. The records owner confirms no legal hold, complaint, or open investigation touches batch B-4402, and QA authorizes destruction | The legal hold check runs before destruction regardless of how routine the record now looks; the destruction record is created and itself retained |
Two things make this trace defensible rather than lucky. First, every handoff in the middle, LIMS to report, CDS to archive, has its own verification step; nothing simply “moved” without a check. Second, the year-6 retrieval works only because the year-1 decommissioning treated archival as a data integrity event and proved retrievability before the original system was switched off. A program that gets generation and review right but skips the archive-verification step is a program that will produce R-5518 perfectly for the first five years and fail to produce it at all in year six, at exactly the moment, a complaint investigation, when it matters most.
On-Premises vs Cloud Archival: Comparative Risk
Where the archive physically lives, on company-owned infrastructure or with a cloud or hosted provider, changes which failure modes need a specific control. Neither option is inherently safer; each concentrates a different risk that has to be planned for rather than assumed away.
| Dimension | On-premises archive | Cloud or hosted archive |
|---|---|---|
| Media and format obsolescence | Company controls the refresh cycle directly, but has to run it; nobody else will notice a drive or tape aging out | Provider typically handles underlying hardware refresh, but the company still owns confirming the logical format stays readable |
| Retrieval dependency | Retrieval depends on internal staff, internal systems, and internal continuity; a site event can disrupt access | Retrieval depends on network access and the provider’s availability and support; a provider outage or contract dispute can block access even when the data is intact |
| Backup and redundancy | The company designs and owns the redundancy; gaps are the company’s own responsibility to find | Often built into the service, but “the provider probably backs this up” is an assumption, not a control, until it is verified in the contract and tested |
| Vendor exit and data portability | Not applicable in the same way; the company always holds the data | A real, planned risk: if the contract lapses, the vendor is acquired, or the service is discontinued, the exit terms and export format decide whether the archive is still usable. Confirm this with an actual test export, not only a contract clause |
| Legal hold and data residency | Under direct company control | Requires confirming the provider can execute and honor a hold, and that data residency terms meet the applicable jurisdiction’s requirements |
| Audit trail preservation across a platform change | Company controls the timing and method of any migration | A cloud platform upgrade or replatforming can change audit trail format or access without the customer initiating it; confirm change notification terms in the contract |
The practical takeaway is not “choose cloud” or “choose on-premises.” Each model has an obsolescence and continuity risk that will not surface until the archive is actually needed, which is precisely the situation you cannot afford to discover a gap in. An on-premises archive needs a documented technology refresh plan. A cloud archive needs a documented, tested exit and export path. Both need the same thing this article has argued for throughout: a periodic, timed retrieval test, run against the real archive, not assumed from a service description or a hardware spec sheet.
Classifying and Controlling Data at Any Stage
Faced with an unfamiliar record type, a team new to lifecycle governance often does not know where to start. The sequence below is a practical gate: work it in order for any data element to land on the right classification and control.
Every terminal box in this sequence points back to a control described earlier in this article: the raw data definition from generation, the hybrid reconciliation from the section above, the migration and archival verification from decommissioning, and the retrieval test from stage six. The gate does not introduce new controls; it tells you which of the controls you already have applies to the record in front of you.
Common Mistakes and Inspection-Finding Patterns
Patterns that recur in inspection findings across the GxPs (generic, not tied to any firm):
- Deleting the dynamic original after printing. Printed results filed, electronic raw data purged for “disk space.” The record can no longer be reconstructed.
- Reprocessing with no prior value or reason. The audit trail shows the new result but not what it replaced or why.
- Audit trails not reviewed at release. Review happens annually or never; the manipulation that the audit trail would have shown ships with the batch.
- Shared or generic logins. Attribution collapses; the audit trail points to “admin” or a team account.
- Backups never restored. A backup exists on paper but has never been proven to restore a usable record.
- Migration that drops metadata. Result values move, audit trail and timestamps do not; the migrated record is no longer attributable.
- Archive that cannot be searched. Data exists on tape but takes days to produce, failing the readily-available requirement.
- Destruction without authorization or under legal hold. Records destroyed early, or while an inspection or investigation is open, with no traceable authorization.
- Selective reporting. A failing result quietly invalidated and omitted from the summary, breaking the completeness principle.
Findings mapped to lifecycle stage
The patterns above recur so often that it helps to see them sorted by where in the lifecycle they occur, because that is usually how an investigator’s questions are sequenced too.
| Lifecycle stage | Recurring finding pattern | Typical citation basis |
|---|---|---|
| Generation | Dynamic original deleted after printing; local, unaudited saves | 21 CFR 211.194; Part 11 |
| Processing | Reintegration or reprocessing with no prior value or reason recorded | 21 CFR 211.194; Part 11 audit trail expectations |
| Review | Audit trails not reviewed before release, or reviewed only annually | FDA 2018 Data Integrity Q&A, routine scheduled review recommendation |
| Reporting | Selective reporting; a failing result invalidated and omitted from the summary | 21 CFR 211.180, 211.194 |
| Retention | Backups never restored; proprietary format no longer readable | 21 CFR 211.180; EU GMP Annex 11 |
| Retrieval | Records exist but cannot be produced promptly, including from a retired system | 21 CFR 211.180(c) |
| Disposition | Destruction with no authorization, or destruction under an open legal hold | 21 CFR 211.180; company retention schedule |
None of these citations is a substitute for reading the guidance directly. They are here so a review of an audit trail or a decommissioning plan can be tied back to the specific stage and the specific expectation it is meant to satisfy, rather than treated as a single undifferentiated “data integrity” bucket.
Interview Questions and How to Answer Them
Questions an inspector or interviewer asks on this topic, with the answers that show you understand the lifecycle.
“What is the original record for this test?” Name the specific file or record that is created first and carries the complete result plus metadata and audit trail, and state whether it is static or dynamic. For a chromatography assay, that is the dynamic raw data file in the data system, not the printout. Point to your written raw data definition.
“What is the difference between a static and a dynamic record, and why does it matter for retention?” A static record is fixed at creation (PDF, printout). A dynamic record can be reprocessed and requeried (a chromatography data file, a LIMS record). It matters because retaining only a static export of a dynamic record loses the metadata and reprocessing capability, so you can no longer verify the result against the raw data. You must retain dynamic records in their dynamic form.
“When does your QA review audit trails?” As part of the batch release or data review flow, not only at annual periodic review, consistent with the 2018 guidance recommendation for routine scheduled review based on system complexity and use. For high-volume systems we use risk-based review by exception, with the sampling tied to a documented risk assessment and full review triggered on defined exception events.
“How do you verify a data transfer between two systems?” Checksum or hash comparison, row-count and sum reconciliation, spot-checks against source, and validated interfaces with error handling for ongoing transfers. The verification is recorded and signed, and interface changes go through change control.
“How long do you retain batch records, and why?” Cite the basis: 21 CFR 211.180(a) for US batch records (expiry plus one year, or three years for certain OTC products without dating), EudraLex Volume 4 Chapter 4 in the EU (one year past expiry and five years past certification, whichever is longer). Note that we standardize on a single longer period defined in procedure to apply it consistently.
“Prove you can retrieve a ten-year-old record from a retired system.” Describe the timed retrieval test: pick a batch, request the complete set including audit trail, time it, confirm completeness against the raw data definition, remediate gaps via CAPA. Tie it to the readily-available requirement in 21 CFR 211.180(c).
“Is metadata part of the record?” Yes. The 2018 guidance defines it as the contextual information required to understand data: user, timestamp, instrument, processing, batch links, audit trail. Complete and Attributable in ALCOA+ cannot be met without it, so it must travel with the data through transfer, archival, and migration.
“Walk me through your disposition process.” Schedule flags end of retention, owner confirms the period expired, legal hold check clears, QA authorizes in writing, destruction is executed, and a destruction record is created and itself retained. Legal hold overrides the routine schedule and the override is auditable.
“How do you keep metadata intact when data moves from a historian or a LIMS into a batch record or ERP system?” Name the authoritative system for each metadata field before the interface goes live, using a documented field-by-field mapping, and verify that mapping when either side changes. Treat a receiving system’s narrower schema as a design risk to check for, not an assumption, and keep the source system as the record of full metadata rather than treating a downstream summary as equivalent.
“How would you approach retiring a system that holds ten years of GxP records?” Build a complete inventory of every record set with a disposition, migrate or archive, before touching anything technical. Verify the migration or archive by reconciliation, not by assuming the export worked. Run an actual retrieval test and confirm records open, are complete, and carry their metadata and audit trail. Only then remove access, and only after that, sanitize media. If the legacy system itself was never properly validated, assess that gap before trusting what comes out of it.
“What’s the difference in archival risk between keeping records on-premises versus in the cloud?” On-premises archives put the technology refresh and continuity burden entirely on the company; the risk is a quietly aging medium or format nobody is tracking. Cloud archives shift day-to-day infrastructure risk to the provider but introduce vendor exit and data portability risk; the archive is only as good as the tested export path if the contract ends. Neither is inherently safer. Both require a periodic, timed retrieval test against the real archive, not a description of the service.
For someone new to GxP, the takeaway is that integrity is not a moment, it is a chain, and the chain breaks most often at the stages nobody is watching. For a working practitioner, the highest payoff is in audit trail review scope, verified transfers, and tested retrieval. For a program-level reader, the deliverable is a system-by-system data flow map with named owners and a retention and disposition schedule that legal hold can override.
The data governance framework article covers the organizational and policy structure for this. Audit trail design and review covers the technical controls that support lifecycle integrity at the generation and processing stages, and data criticality and data risk helps you decide where to spend control effort first. Together with this lifecycle view, they make up the spine of a data integrity program that holds up under inspection.