This is a ready-to-use organizational readiness assessment. It answers a different question than a model validation package: not “does the model perform,” but “is the organization around the model ready to supervise it.” Run it before any AI use case goes live in a GxP process, and again on a defined cadence for models already in production. Replace every <<FILL: ...>> placeholder, score every dimension against evidence rather than opinion, and route the completed assessment through your normal quality and AI-governance channels. A worked filled specimen follows the template. This is an educational aid to adapt and verify against your own quality system, not a compliance guarantee.
Document control header
| Field | Entry |
|---|---|
| Document title | AI Workforce and Organizational Readiness Assessment |
| Document number | <<FILL: CHK-ID, e.g. CHK-QA-041>> |
| Version | <<FILL: version, e.g. 1.0>> |
| Assessment date | <<FILL: date>> |
| AI use case / model assessed | <<FILL: name and register ID>> |
| Assessed by | <<FILL: name, role>> |
| Reviewed by (AI steward / QA) | <<FILL: name, role>> |
| Applies to | <<FILL: site / function / use case in scope>> |
How to use this assessment
- Score each dimension Ready, Partial, or Not ready against the evidence prompt in that row, never against a general impression of how the program is going.
- A rating with no evidence entry is not a rating; leave it Not ready until evidence exists.
- Do not average the scores into a single headline number. A single Not ready row can be the entire reason a deployment is not defensible, regardless of how well the others score.
- Use this per use case as a go-live gate, and re-run it on a fixed cadence, and after any material change to the model, the team, or a related incident, for models already in production.
- File the completed assessment as a governance record. It is itself evidence that readiness was assessed on purpose rather than assumed.
Readiness dimensions
| # | Dimension | Evidence prompt | Rating | Evidence |
|---|---|---|---|---|
| 1 | Data literacy | Can the people who will supervise this model’s output correctly interpret a confidence score, a false positive, and a false negative for this specific use case, demonstrated on a scenario, not asserted? | <<FILL: Ready / Partial / Not ready>> | <<FILL>> |
| 2 | Roles filled | Is a named model owner and a named AI steward assigned to this use case, distinct from the people who built it? | <<FILL>> | <<FILL>> |
| 3 | Citizen development governance | If any part of this use case was built or configured by staff outside a formal data-science or IT function, is it inventoried, risk-tiered, and paired with a steward? | <<FILL>> | <<FILL>> |
| 4 | Training | Are the reviewers or supervisors for this model trained specifically on its known failure modes, with a recorded applied-judgment assessment rather than a read-and-understand signature? | <<FILL>> | <<FILL>> |
| 5 | Staffing pattern decided | Has the organization deliberately decided and documented whether this use case is human-in-the-loop or human-on-the-loop, and staffed to that decision rather than to convenience? | <<FILL>> | <<FILL>> |
| 6 | Human-AI partnership health | Is the human genuinely engaged rather than rubber-stamping: is an override or disagreement rate tracked, and does it sit at a level consistent with the model’s known error rate? | <<FILL>> | <<FILL>> |
| 7 | Operating model | Does this use case have an end-to-end lifecycle owner, with every handoff (build to validation, validation to operations, operations to monitoring, monitoring to retrain) named and evidenced? | <<FILL>> | <<FILL>> |
| 8 | Monitoring ownership | Is a named person accountable for production monitoring, with a defined trigger-to-response path that has actually been exercised or tested? | <<FILL>> | <<FILL>> |
| 9 | Change management | Is the model used as designed with no documented workaround, and can affected staff articulate both what it is good for and where it fails? | <<FILL>> | <<FILL>> |
| 10 | Governance culture | Can staff describe, credibly, a time they raised a concern about this model or a similar one and what happened, and is there no throughput incentive that structurally rewards over-trusting the output? | <<FILL>> | <<FILL>> |
Scoring and the go-live decision
| Result pattern | Recommended action |
|---|---|
| All ten dimensions Ready | Proceed to or continue production use at the intended scope. |
| One or more Partial, none Not ready | Proceed only with the Partial dimensions carrying a dated, owned closure action; re-check at the next milestone or within a defined interval. |
| One or more Not ready | Do not deploy at the intended scope. Either close the Not ready gaps first, or scope the deployment down to a lower-risk pattern (for example advisory-only, human-in-the-loop, a limited pilot population) that the organization can currently support, and re-assess before expanding. |
Acceptance criteria
This assessment is complete and usable when all of the following are true:
- Every dimension carries a rating and a specific evidence entry, not an assertion.
- No Not ready dimension is deployed around without either closing the gap or formally scoping the deployment down, with that decision recorded and approved.
- The assessment is signed by the assessor and reviewed by the AI steward or QA before the go-live decision is made.
- The assessment is re-run on the defined cadence for any model already in production, not only before first deployment.
References
21 CFR 211.25(a) (personnel education, training, and experience, including training on a continuing basis), and EU GMP EudraLex Volume 4, Part I, Chapter 2 (personnel and training), as the basis for treating supervisory competence over an AI system as a training requirement subject to inspection. FDA guidance, “Data Integrity and Compliance With Drug CGMP: Questions and Answers” (December 2018), and MHRA “GXP Data Integrity Guidance and Definitions,” on management responsibility for a work environment and culture that supports reliable records, extended here to AI-supervising behavior. PIC/S PI 041, Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments, on data governance as an organizational, not purely technical, control. ICH Q9(R1), Quality Risk Management, as the basis for scaling organizational readiness effort to the risk tier of the use case. ICH Q10, Pharmaceutical Quality System, on management review as the forum where readiness gaps and their closure are reported.
Confirm the current version of each reference before you rely on it.
Revision history
| Version | Date | Author | Summary of change |
|---|---|---|---|
<<FILL: 1.0>> | <<FILL: date>> | <<FILL: author>> | Initial issue. |
Approvals
| Role | Name | Signature | Date |
|---|---|---|---|
| Assessor | <<FILL>> | ||
| AI steward | <<FILL>> | ||
| Quality approver | <<FILL>> |
Filled specimen
The following shows the assessment completed for an illustrative QC data review AI-assist tool that flags results at elevated risk of being out of specification before final analyst review. The organization, ratings, and evidence are illustrative; replace them with your own.
| # | Dimension | Rating | Evidence |
|---|---|---|---|
| 1 | Data literacy | Partial | Senior analysts correctly interpreted the confidence output in a five-scenario test; two of six junior analysts on the same rotation could not distinguish a low-confidence flag from a high-confidence one |
| 2 | Roles filled | Not ready | No AI steward named for the QC domain; the data scientist who built the model was the informal point of contact, with no quality-side counterpart |
| 3 | Citizen development governance | N/A | Built entirely by the internal data-science team on the sanctioned MLOps platform; no citizen-developer component |
| 4 | Training | Partial | Training records existed but were read-and-understand only; no case-based judgment assessment and no coverage of the model’s known low-end blind spot on one assay |
| 5 | Staffing pattern decided | Ready | Documented and approved as human-in-the-loop: the model flags, the analyst decides every case, no per-output action is taken automatically |
| 6 | Human-AI partnership health | Not ready | No override-rate tracking existed; the six-week pilot had no data on how often analysts disagreed with a flag |
| 7 | Operating model | Partial | Build-to-validation handoff was documented; nothing was defined for after go-live |
| 8 | Monitoring ownership | Not ready | No named monitoring owner and no defined drift trigger or response path |
| 9 | Change management | Partial | Pilot-team analysts were briefed and supportive; analysts outside the pilot team who would inherit the tool at wider rollout had not been told anything |
| 10 | Governance culture | Partial | No throughput pressure identified, a genuine strength, but no defined channel existed for raising a concern about the model specifically, separate from the general deviation process |
Reading this specimen the way an assessor should: three dimensions scored Not ready (roles, partnership health, monitoring ownership), which blocks full-scale deployment under the scoring rule above. The site’s decision, recorded against this assessment, was to go live in a tightly bounded advisory mode limited to the trained pilot-team analysts, name an AI steward within four weeks, rebuild training around real misclassified cases with a scored assessment, and stand up override-rate tracking with a named monitoring owner before expanding. That decision, and the twelve-week re-assessment that confirmed the gaps closed, is the record an inspector or an auditor would expect to see behind a “we assessed readiness” claim.
Common inspection findings this assessment prevents
- A technically validated model deployed with no organizational readiness assessment on file, so people-side gaps surface only when an inspector finds them.
- A production model with no named steward or owner, so no one can answer who is accountable when something drifts.
- Training records that show attendance but no evidence a reviewer can recognize a model error.
- A human review step with no override-rate data, so rubber-stamping cannot be ruled out or confirmed either way.
- A readiness claim made once at deployment and never revisited, while the model, the team, and the surrounding process all continued to change.
How to adapt this assessment
- Set your document number, the use case being assessed, and the assessor and reviewer names in the header.
- Adjust the evidence prompts to name your real systems, roles, and terminology, but keep every prompt evidence-based rather than a yes/no on intent.
- Set your re-assessment cadence and the specific triggers (model retrain, team change, related incident) that force an out-of-cycle re-run.
- Route the go-live decision through your existing AI governance structure so the assessment feeds a real gate, not a filed-and-forgotten form.
- Confirm every reference against its current published version before issue.