Most organisations that commission an AI readiness assessment already believe they know the answer. They have run pilots, hired data scientists and signed off a strategy deck, and they expect the assessment to confirm that the hard part is behind them. In my experience it rarely is. The gap between feeling ready for AI and being ready to depend on it is one of the most consistent, and most expensive, misjudgements I see in executive teams.
To make the pattern concrete, let me describe an organisation. It is a composite, drawn from several engagements rather than a single client, but every detail below is something I have encountered first-hand.
Picture a mid-sized insurance intermediary. The CEO has commissioned an AI strategy from a well-known consultancy, and the resulting deck places the firm comfortably on the upward slope of a maturity curve. There is a data lake. There is a head of data science, hired eighteen months earlier. There have been two proofs of concept, one for claims triage and one for a customer-facing assistant, and both produced demo videos the leadership team has watched more than once. The brief to me is to “scale it across the business”, which is the phrase people use when they believe the difficult work is done.
It takes about twenty minutes to see that it is not. The claims triage model has never been connected to the claims system; it was trained on an exported spreadsheet that a data scientist refreshes by hand every few weeks. The customer assistant was switched off after giving three policyholders incorrect information about their excess, and nobody in the room can say who made that decision, when, or on whose authority. The data lake turns out to be a place where several systems dump files every night, with no owner, no schema governance and no one who can vouch for whether a given field is reliable. The organisation has the appearance of readiness. What it lacks is the structural foundation that lets AI survive contact with a real business.
That gap matters because organisations that misjudge it do not fail cheaply. They fail after committing budget, restructuring teams and making promises to boards and customers. An AI readiness assessment done honestly and early is the cheapest intervention available to an executive. It is also the one most often skipped, because the questions are uncomfortable and the answers are rarely the ones leadership hoped for.
Why Maturity Curves Mislead
The standard tool for assessing AI readiness is the maturity model: a diagram with four or five stages, from “experimenting” to “transformative”, on which leaders locate themselves. Its appeal is obvious. It turns a complex question into a single point on a line, and it tends to place the organisation paying for it somewhere flattering. I have yet to see a consultancy maturity assessment that told a client they were at stage one.
Optimism is the smaller problem. The bigger one is that maturity models measure activity: how many pilots you have run, whether you have hired data scientists, whether a strategy document exists. They then treat that activity as evidence of readiness. Running two pilots makes you experienced at running pilots, which is a much smaller achievement than being ready. A head of data science with no authority over the systems their models must connect to does not make you ready either.
Readiness is a measure of whether the organisation has the preconditions for AI to create value without creating unacceptable risk. Those preconditions are data ownership, decision rights, operational integration and a real willingness to change how work gets done. They do not show up well on a maturity curve, because they are not milestones you pass once. They are properties of the organisation, and most organisations do not yet have them.
There is a second confusion worth naming, and it is the one this audit is built around: the difference between being ready to experiment with AI and being ready to depend on it. Almost any organisation is ready to experiment. Experiments are contained, low-stakes and forgiving. Dependence is the opposite. Once an AI system sits in a live process that customers or revenue rely on, the tolerances change completely. The questions that matter for dependence (who is accountable when the model is wrong, how you detect degradation, what the fallback is) are precisely the ones experimentation lets you ignore. Most readiness failures are really a failure to notice that you have crossed from one mode into the other.
The AI Readiness Assessment: Twelve Questions in Four Domains
The diagnostic I use has twelve questions across four domains. I use it both with organisations that believe they are ready and with those about to commit serious budget. It is deliberately not a maturity curve. Its job is to expose the specific gaps that will decide whether an AI investment compounds or collapses.
Two rules before you start. First, score a specific use case, not the whole organisation. Readiness is contextual: you may be ready in one function and nowhere near it in another. Second, score on evidence, not intention. Each question earns two points if you can answer with concrete, verifiable evidence, one point for a partial or aspirational answer, and zero if the honest answer is “we don’t know” or “we assumed someone else had this”.
Domain one: Data foundations
1. Do you know who owns the data behind this use case, and is it accessible in production rather than as an export? Our composite insurer would score zero here: its triage model ran on a hand-refreshed spreadsheet.
2. Do you know how good that data is (its completeness, accuracy and freshness), measured rather than assumed? Data quality is the most common silent killer of AI initiatives. Models trained on plausible-looking but flawed data produce plausible-looking but flawed outputs, and the flaws surface only in production.
3. Can you trace where the data came from, and on what legal basis it may be used for this purpose? Lineage and permissions are the difference between a system you can defend to a regulator and one you cannot. Under GDPR and UK GDPR, data collected for one purpose cannot simply be repurposed to train or run a model.
Domain two: Decision rights, accountability and regulation
4. When the system produces an output, who is accountable for the decision that follows: a named person, not a committee? If accountability is distributed or unclear, you are ready to experiment with the system, not to depend on it.
5. Who has the authority to switch the system off, and is that authority documented and known? In the composite case, the fact that nobody could say who had disabled the customer assistant was itself the finding.
6. Have you classified this use case against the regulation that applies to it, and defined what the system may do autonomously versus what requires a human decision? For organisations operating in the EU, that means knowing where the use case sits in the AI Act’s risk tiers. The obligations for high-risk systems, including human oversight, are substantial, and some sectoral uses, such as risk assessment and pricing in life and health insurance, are explicitly listed. In the UK, the equivalent questions arrive through UK GDPR’s rules on automated decision-making and through sector regulators; in financial services, the FCA’s Consumer Duty. An organisation that cannot answer this has been treating the system as a feature, when it is really an actor in its operating model.
Domain three: Operational integration
7. Is the system connected to the real systems of record, or does it run on a copy? A model that changes nothing in the actual workflow is a demo, however impressive.
8. When the model is wrong (and it will be), what is the fallback, and has it been tested? An untested fallback is the signature of an organisation that has mistaken a pilot for a production system.
9. How will you detect when performance degrades? Models drift as the world they were trained on changes. Without monitoring, you find out through customer complaints, which is the most expensive detection method there is.
Domain four: Organisational capacity
10. Are the people whose work will change prepared for it, and were they involved in designing the change? AI initiatives fail more often through quiet operational rejection than through technical inadequacy.
11. Can you maintain and improve the system internally, or does it depend entirely on a vendor or a single individual? Dependence on one person’s knowledge is a risk that grows unnoticed until that person leaves. It also has a regulatory side: the EU AI Act already requires organisations deploying AI to ensure a sufficient level of AI literacy among the staff who operate it.
12. Are you prepared to stop if the evidence says the value is not there? This is the question that separates serious organisations from the rest. A leadership team that cannot say no to its own AI initiative has not decided to adopt AI. It has made a commitment it can no longer evaluate.
The Scorecard
Use the table below to score each question for a single, named use case. If you cannot confidently place an answer in the “2” column, it is not a 2.
#
Question
Scores 2 when…
Scores 0 when…
1
Data ownership and production access
A named owner exists and the model reads live production data
Data arrives by manual export, or nobody owns it
2
Measured data quality
Completeness, accuracy and freshness are measured and reported
Quality is assumed because “the data looks fine”
3
Lineage and legal basis
Sources are traceable and the lawful basis for this use is documented
Nobody can say where key fields come from or whether use is permitted
4
Named accountability
One named person owns decisions based on the output
Accountability sits with “the project” or a committee
5
Authority to switch off
A documented kill-switch owner and trigger conditions exist
Nobody knows who can, or did, turn it off
6
Regulatory classification and autonomy limits
Risk tier is assessed and autonomous vs human-approved actions are written down
Classification has not been considered and limits are implicit
7
Integration with systems of record
Outputs flow into the live workflow
The model runs on a copy and outputs are read, not used
8
Tested fallback
A fallback process exists and has been exercised
“We’d just turn it off”, untested
9
Degradation monitoring
Performance is tracked against outcomes with alert thresholds
Problems would surface via complaints
10
People involved in the change
Affected staff helped design the new process
Staff will be told when it goes live
11
Internal capability
At least two internal people can maintain it; operators are trained
One person or one vendor holds all the knowledge
12
Willingness to stop
Stop criteria are agreed in advance
Stopping is not discussed, or not politically possible
Reading your score (out of 24):
0–11: Not ready to depend on AI in any process that matters, however many pilots you have run. Keep experimenting, and close the gaps.
12–18: Ready for bounded production use with strong human oversight and a narrow scope.
19–24: Ready to scale with confidence. This is rare.
The total matters less than where the zeros are. A single zero on questions 4, 5, 6 or 8 (accountability, switch-off authority, regulatory classification or fallback) should stop a production deployment on its own, whatever the total.
How the Recovery Usually Goes
Back to our composite insurer. When leadership teams in this position run the audit properly, the score is usually sobering; a result around nine is typical. The room goes quiet, because the number contradicts eighteen months of internal narrative and an expensive consultancy deck. The organisations that go on to succeed are the ones whose CEO does not argue with the number but asks what to do about it.
The first move is to stop talking about “scaling AI across the business” and pick a single use case whose gaps are fixable. Here, that means internal claims triage rather than the customer-facing assistant. Ownership of the claims data goes to a named person in operations, with the authority to enforce a schema. The model is connected to the live claims system rather than a nightly export, which typically takes a couple of engineers most of a quarter. The limits are written down: the model may route and prioritise claims but may not decline or settle them, and it is documented who can switch it off and under what conditions. A monitoring dashboard tracks routing decisions against outcomes, so that degradation shows up as a metric rather than a complaint.
What works is narrowing the scope. Concentrating on making one thing production-ready, instead of attempting everything, produces a system that within months handles most routine triage with measurable reductions in cycle time. What fails, at least at first, is the human side. When the claims handlers whose work is being reshaped have not been involved early, they route around the system, exactly as question ten predicts. The fix is to go back, involve them in refining the routing logic and redeploy with their input. The lesson is that organisational capacity cannot be bolted on afterwards.
Where This Audit Reaches Its Limits
A readiness audit sold as a silver bullet would betray the discipline it is meant to enforce, so it is worth being clear about what this one does not do.
It is a diagnostic, not a strategy. It will show you where your gaps are, but it will not tell you whether AI is the right investment for your business in the first place. That is a prior question of strategy and economics. An organisation can score 24 and still be pursuing an initiative that solves a problem not worth solving.
It can also be turned into an excuse for inertia. In cautious organisations, a low score can become permission to do nothing indefinitely, which is its own form of failure. The right response to a score of nine is to fix the specific gaps the audit identified, not to abandon AI. Telling prudent preparation apart from disguised avoidance takes judgement that no scoring system supplies.
It is not a compliance assessment. Question six asks whether you have classified your use case; it does not do the classification for you. Where the answer points to high-risk obligations, you need proper legal and technical review.
Finally, it assumes honest answers. Its worst failure mode is a leadership team that scores itself on aspiration rather than evidence. An audit filled in with the answers you wish were true is worse than no audit, because it dresses false confidence up as rigour.
Replace the Feeling With Evidence
The twelve questions matter, but the deeper point is what “readiness” is taken to mean. For most executives, readiness is a feeling produced by activity: pilots run, people hired, strategies commissioned. That feeling is unreliable because it measures effort, not preparedness. The audit’s real job is to replace the feeling with evidence, and to force a clear distinction between being busy with AI and being ready to depend on it.
Across the organisations that handle this well, the strongest predictor of success I have found is a willingness to score yourself low. The ones that fail rarely lack capability. They are the ones that could not tolerate an honest answer to a simple question. AI does not reward the organisations that want it most. It rewards those prepared to hear that they are not ready, and to fix that before they build anything at all.
Run the audit on your own use case. Check your AI Readiness Scorecard — a one-page version of the 6 questions and scoring criteria to use with your leadership team.


