This is a ready-to-use periodic monitoring log for an AI screening tool that is already live. Two numbers tell you whether it still works and whether the human checkpoint is real: the reviewer override rate, and the tool’s score on a small held-out set of known cases. A rising override rate warns of drift; a near-zero rate warns of rubber-stamping; a degrading held-out score is your objective early warning. This log tracks both on a schedule and ties them to a documented off-switch. Replace every <<FILL: ...>> placeholder and run it under the tool’s parent SOP. A filled specimen follows. This content is educational reference, not legal or regulatory advice; adapt it to your validated tool and quality system.
Header
| Field | Entry |
|---|---|
| Log / form number | <<FILL: FORM-ID>> |
| Tool / system name and ID | <<FILL>> |
| Governing SOP | <<FILL: SOP-ID>> |
| Monitoring frequency | <<FILL: e.g. monthly>> |
| Pinned model / prompt version | <<FILL>> |
| Metric owner | <<FILL: QA role>> |
How to read the override rate
The override rate is the share of the tool’s flags (or suggestions) that the human reviewer changed. Interpret it against these bands, and act on a shift, not a single reading.
| Override rate | Interpretation | Action |
|---|---|---|
| Under 10% | Model and inputs stable, reviewers engaged | Continue, monitor |
| 10 to 25% | Some drift or a shifted input mix | Investigate the overridden categories for a pattern |
| Over 25% | Model drift, prompt staleness, or an unseen process change | Trigger periodic review early; consider pausing the suggestion |
| Near 0% over a long window | Likely rubber-stamping, not real review | Audit the review step; the checkpoint may be theater |
Both extremes matter to an inspector: a high rate signals the tool is degrading, a near-zero rate signals nobody is really reviewing.
The held-out set
Keep a small, fixed set of known-good and known-bad cases with the correct answer recorded, and run the tool against it each period. This is the objective check that does not depend on reviewer behavior. Measure recall on the known-bad cases especially: a tool that starts missing a category of real problem is degrading whatever the override rate says.
Monitoring grid (blank)
| Period | Flags / suggestions | Overrides | Override rate | Held-out recall (known-bad) | Held-out precision | Drift signal? | Action taken | Reviewer | Date |
|---|---|---|---|---|---|---|---|---|---|
<<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL>> | <<FILL: Y/N>> | <<FILL>> | <<FILL>> | <<FILL>> |
Off-switch triggers
Record and act on any of these; each is a documented reason to pause or stop the tool and notify the system owner:
- Override rate over
<<FILL: threshold, e.g. 25%>>for<<FILL: e.g. one period>>. - Held-out recall on known-bad cases below
<<FILL: threshold>>. - A model or prompt change not yet assessed under change control.
- A process or input change the tool has not been validated against.
- Any period with a near-zero override rate that a review-step audit cannot explain as genuine.
Acceptance criteria
- The override rate and the held-out results are recorded every period at the defined frequency.
- Any band shift or off-switch trigger has a recorded action, not just a number.
- The held-out set is fixed, with correct answers on record, and run each period.
- A triggered off-switch results in a pause or stop and a notification, tracked to closure.
- The log is signed and retained per
<<FILL: retention>>.
References
FDA guidance, Computer Software Assurance for Production and Quality Management System Software. 21 CFR Part 11 and the predicate rule for the activity. FDA and EMA, Guiding Principles of Good AI Practice in Drug Development (January 2026), for lifecycle management. Quality metrics practice under ICH Q10.
Confirm the current version of each reference before use.
Filled specimen
Three monitoring periods for a deviation-categorization suggestion tool. Illustrative only.
| Period | Flags / suggestions | Overrides | Override rate | Held-out recall (known-bad) | Held-out precision | Drift signal? | Action taken | Reviewer | Date |
|---|---|---|---|---|---|---|---|---|---|
| Apr 2026 | 210 | 16 | 7.6% | 0.96 | 0.93 | N | Continue | J. Okoro | 03 May 2026 |
| May 2026 | 198 | 41 | 20.7% | 0.90 | 0.88 | Y | Investigated overrides: a new deviation type the rubric did not cover; rubric update proposed under change control | J. Okoro | 02 Jun 2026 |
| Jun 2026 | 205 | 12 | 5.9% | 0.95 | 0.92 | N | Rubric update effective; rate back in band | J. Okoro | 03 Jul 2026 |
The May spike into the 10 to 25% band was not ignored; the owner investigated the overridden cases, found a new deviation type the classification rubric had never seen, and updated the rubric under change control. June confirms the fix held. The held-out recall stayed high throughout, so the tool was never blind to known problems, and the whole episode is documented as active lifecycle management rather than a silent drift.
Common inspection findings this log prevents
- A live AI tool with no evidence anyone monitors whether it still works.
- Override rates near zero for months, showing the human review is a rubber stamp, with no audit.
- Model drift that reached records because degradation was never measured against a held-out set.
- A model or prompt updated to “improve accuracy” with no assessment and no monitoring of the effect.
- No documented condition under which the tool would be paused or stopped.
How to adapt this log
- Set your form number, governing SOP, frequency, and thresholds in the header and off-switch section.
- Build and freeze a held-out set of known cases with correct answers, sized to your workflow.
- Compute the override rate from your disposition log each period and record both metrics here.
- Make every band shift or trigger an action with an owner, tracked to closure.
- Confirm every reference against the current published version before use.