Most companies can't tell if their AI is actually paying off

Most companies are not failing to get value from AI. They are failing to find out whether they did. Here is why returns go missing, and what closes the gap.

Most companies are not failing to get value from AI. They are failing to find out whether they did.

McKinsey's global survey of 1,933 organisations found that 88% regularly use AI in at least one business function. Only about 6% attribute more than 5% of EBIT to it. Around 39% report any enterprise-level EBIT impact at all, and most of those put it under 5%.

Read those two numbers together and the story is not "AI does not work". Nearly nine in ten organisations kept using it. People do not keep using tools that do nothing. The story is that something real is happening and almost nobody can prove it, price it, or repeat it deliberately.

That is a different problem, with a different fix.

Reason one: nobody measured

Ask a company what its AI programme returned and you usually get one of three answers. A confident number with nothing behind it. An honest shrug. Or a deflection to activity metrics: seats deployed, users onboarded, prompts run.

The reason is almost always the same. Nobody captured a baseline.

Return on investment needs a before and an after on the same metric, and ideally a group who did not get the tool, so you can tell the difference between the AI and everything else that changed that quarter. Very few rollouts do this. The pilot starts because it is obviously a good idea, and the question of proof arrives nine months later when somebody senior asks what the £400k bought.

At that point it is too late. A baseline reconstructed after the fact is not a measurement, it is a rationalisation. You are asking people to remember how long something used to take, and they will remember it in whatever way makes the current situation make sense.

This is how folklore fills the vacuum. The most quoted statistic in enterprise AI right now is that 95% of AI pilots fail. It comes from a preliminary working paper, not peer-reviewed, written by a group that builds agentic AI infrastructure. The 95% refers to organisations reporting zero return, not to 95% of pilots failing, and reconstructing from the paper's own funnel gets you closer to 75% of pilots missing a deliberately strict bar. It is repeated in board papers every week as though it were a law of physics.

A number that shaky travels that far for one reason: there is nothing better to put in its place.

Reason two: the ones who do measure count the wrong things

Suppose a company does try. Here is what usually gets counted, in roughly this order.

Seats deployed. This measures procurement, not value. It is the number most confidently reported to boards and the one that means least.

Active users. Better, but it tells you people opened the tool, not that anything changed as a result.

Self-reported hours saved. This is where most programmes land, and it is worth being blunt about how weak it is.

In a randomised controlled trial, METR gave 16 experienced open-source developers 246 real tasks from their own repositories, randomly allowing or forbidding AI tools on each one. The developers expected AI to make them 24% faster. Measured, they were 19% slower. Afterwards, having lived through it, they still believed AI had made them 20% faster.

METR is careful about how far this generalises, and so should we be: one setting, one snapshot of early-2025 tooling, 16 developers. The dramatic reading, that AI slows people down, is not what the study supports. The narrow reading is better evidenced and far more useful to anyone building a business case: people cannot feel their own productivity. The error here was large, and living through it did not correct it.

If self-perception is that unreliable in a group of expert engineers doing measurable work, a survey asking marketing how many hours AI saved them last month is not data.

Then there is the arithmetic that follows. It usually looks like this:

200 people × 5 hours saved per week × 46 weeks × £65 per hour = £3m of value created

Every AI business case in the country contains a version of this sum. It falls over on one question, which a good CFO asks immediately: which cost line went down?

Usually none of them. The hours were saved and then absorbed. People used them to do more of the same work, or to attend more meetings, or simply to have a less compressed day. That may be genuinely worthwhile. It is not £3m, and presenting it as £3m is how AI budgets lose credibility.

Saved time only becomes value at the point it is redeployed into something that shows up. Almost nobody tracks redeployment, which is why productivity metrics look strong while finance sees nothing.

The one worth counting instead: what does it now cost to get one finished, correct piece of work, end to end, including the human time spent checking and fixing it?

That last clause matters more than it sounds. Research from BetterUp Labs and Stanford's Social Media Lab, published in Harvard Business Review, named the thing everyone had noticed but not measured: workslop, AI output that looks finished but is not. In a survey of 1,150 US workers, 41% had received some in the previous month, and each instance took an average of one hour and 56 minutes to sort out. For the workers on the receiving end that came to around $186 a month each, which scales to over $9m a year in a 10,000-person company.

None of which is an anti-AI story. It is a measurement story. Time to first draft fell; time to finished work often did not, because the difference landed on somebody else's desk. And that is the crux of it. The person who felt the saving and the person who paid for it are rarely the same person, so neither of them can give you an accurate number on their own.

Reason three: they treated it as a tools and training problem

This is the one I think most companies get wrong, and it gets discussed least.

The standard AI programme has two moves. Buy licences. Run training. Both are sensible, neither is sufficient, and the reason why is easier to see if you lay out the rungs between spending money and getting some back.

RungWhat it meansTypically measured?
AccessPeople have the toolYes, obsessively
AdoptionPeople open the toolYes
ApplicationPeople use it on the work that actually matters, repeatablyRarely
ValueThe result shows up in a number finance recognisesAlmost never

Licences buy you rung one. Training buys you rung two, and some of rung three if it is very good. Nothing in the standard programme carries an organisation from three to four, because nothing in it answers the only question that matters to the person doing the job: what exactly should I use this for, on my systems, in my role, in a way that works the same way every time?

General capability does not answer that. Telling a claims handler that AI is powerful and here is a prompting course is like handing someone a workshop full of tools and a safety briefing, then being surprised that no furniture appears. They have access. They have skill. What they do not have is the plan, the jig, and the worked example.

The difference is concrete. Compare what a company typically gives someone:

"Claude is now available to everyone. Here is a two-hour course on effective prompting, and a link to a shared document of useful prompts."

With what they actually need:

A defined Capability called Draft a liability assessment. It knows your policy wording and your escalation thresholds, pulls the claim record from the claims system, produces the assessment in your house format, flags anything above £25,000 for a human, and logs every run.

The first is a resource. Somebody has to be motivated enough to go and find it, adapt it, and remember it next Tuesday. The second is a piece of working equipment that produces the same output whoever pulls the handle, and, because every run is logged, it produces a measurable one. Only the second thing can ever appear on a P&L, because only the second thing is repeatable enough to count.

And here is the compounding problem: in most organisations, somebody has already figured it out. There is a person in that claims team who has worked out a genuinely excellent way to use AI on a specific recurring task. That knowledge currently lives in their personal chat history. It reaches nobody. It is not reviewed, not improved, not measured, and when they leave it goes with them.

So the same organisation simultaneously has undiscovered value it cannot see, and a training programme teaching everyone else from scratch.

What the companies getting returns do differently

Nothing exotic. Four things, in order.

  1. They find out what is actually being used. Not the sanctioned pilot, everything, including the personal accounts. You cannot compute a return on activity you cannot see, and the most valuable use is frequently the least visible.
  2. They pick named use cases with a real before-number. One process, one metric, captured before anything changes. "Reduce time from claim received to decision issued, currently 4.2 days." Not "improve productivity".
  3. They turn what works into something repeatable. The good approach one person invented becomes a defined, sanctioned Capability the whole team can run the same way, rather than a story told at a lunch and learn.
  4. They measure the finished unit, and the redeployment. Cost per completed piece of work including rework, and an honest answer about what the freed time became.

Notice what step one actually is. Discovery, not measurement, and it comes first. This is the part that gets skipped. Almost every published framework opens at step two, which is unhelpful, because step two is impossible until you know what your organisation is already doing.

The uncomfortable version

The 6% figure is usually read as a story about AI underdelivering. Read it again with the measurement problem in mind and it says something less comfortable.

We do not actually know how much value AI is creating in most companies, because most companies are not instrumented to find out. Some of the 94% are genuinely getting nothing. Some are getting a great deal and cannot see it, because it is happening in personal accounts, on unnamed tasks, invented by individuals, recorded nowhere.

Both groups get the same result at the board meeting, and both will make the same decision next year, which is to cut the budget.

That is what makes this worth fixing properly rather than arguing about. Connor starts at step one for exactly this reason: find the AI already connected to your systems and the people already using it well, turn the good patterns into Capabilities the team can actually run, and keep the record that lets you say what it did.

You cannot manage a return you never measured. You certainly cannot repeat it.

Frequently asked questions

What percentage of companies actually get a financial return from AI?
In McKinsey's global State of AI survey of 1,933 organisations, 88% said they regularly use AI in at least one business function, but only around 6% attributed more than 5% of EBIT to it. About 39% reported any enterprise-level EBIT impact at all, and most of those put it below 5%. Adoption is close to universal. Measurable financial return is rare.
Why is AI ROI so hard to measure?
Mostly because the measurement was never set up. ROI needs a baseline captured before deployment on the same metric you will report afterwards, and ideally a group that did not get the tool for comparison. Very few rollouts do either. Without a baseline there is no counterfactual, so any number produced later is an estimate dressed as a measurement.
Are hours saved a good measure of AI value?
No, for two reasons. Self-reports are unreliable: a randomised trial found developers were 19% slower with AI while believing they were 20% faster. And saved time only becomes value if it is redeployed into work that shows up somewhere. An hour saved and absorbed into the day is invisible to the P&L, however real it felt.
Is getting value from AI a training problem?
Training helps, but it is not sufficient. Licences give people access and training gives them general skill. Neither tells anyone what specifically to do with AI in their own role, on their own systems, in a way that is repeatable and can be observed. That gap between capability and application is where most AI programmes quietly stall.
James ZhaoCo-founder, Connor

James is the co-founder Connor. After a corporate career at Barclays and KPMG as a software engineer, he built and exited his own software company. He has spent the last three years at the forefront of AI, and the most recent of them building AI-native products and the agent platform behind Connor.

All posts

Find out what your team has already built.