The AI projects that clear a board are the ones that touch pipeline, margin or risk through a metric the board already tracks. The ones that stall are overwhelmingly individual-productivity deployments, copilots and assistants bought seat by seat, because nobody can say what happened to the time they saved. That is not a cynical reading; it is what the published evidence now shows. MIT's NANDA initiative reported in August 2025 that about 95 per cent of enterprise generative AI pilots deliver little or no measurable profit-and-loss impact, while roughly 5 per cent achieve rapid revenue acceleration. The Department for Science, Innovation and Technology found in research updated in February 2026 that most UK businesses using AI report a productivity increase, yet most have seen no change in revenue. Productivity that never becomes revenue, cost or reduced risk is exactly the gap this article is about, and it is a gap with a cause that can be fixed.
The argument in brief, drawn from a June 2026 discussion among technology leaders on r/CIO and tested against published research below: individual-productivity AI accelerates existing work, and accelerated work is invisible to a board unless someone deliberately reassigns the time it frees. Process-level AI redesigns the work itself, which forces the reassignment to happen and makes the return legible. The distinction sounds subtle. Financially, it is the whole game.
Why individual-productivity AI keeps failing the board test
Individual-productivity AI fails the board test because organisations measure deployment and usage, not outcomes. A CIO can report that 4,000 copilot licences are active and adoption is at 70 per cent, and none of that is a business result. The question a finance director will ask, and the one almost nobody can answer, is the one posed in that r/CIO thread: suppose the copilot works, and a task that took a week now takes a day. What happened to the other four days? Were they reassigned to something more valuable, absorbed into meetings and low-impact work, or nothing at all? If the organisation cannot answer, the benefit cannot be booked.
The best field evidence available says the honest answer is usually "not much". Economists Anders Humlum and Emilie Vestergaard linked two large adoption surveys, covering roughly 25,000 workers across 7,000 workplaces in 11 exposed occupations, to Danish administrative records on earnings and hours. Their July 2025 working paper found average time savings from AI chatbots of about 3 per cent of working hours, and precisely estimated null effects on earnings and hours, with confidence intervals ruling out effects above 1 per cent. This in occupations, from software development to legal work, where randomised controlled trials had shown productivity gains of 15 to 50 per cent. The tools work in the lab. In the wild, the saved time dissolves, partly into new tasks the tools themselves create, which the authors describe as integration and oversight work.
The Danish study also found that 38 per cent of the surveyed workplaces had deployed enterprise chatbot versions and 30 per cent of employees had received training, so this is not a story about immature adoption. Employers invested; the investment raised usage; usage did not move the numbers a board reads.
UK survey data rhymes with this. The Office for National Statistics reported on 20 July 2026 that self-reported AI use among UK businesses with ten or more employees rose from around 12 per cent to around 35 per cent since late 2023, but describes adoption as shallow: the average number of AI technologies per adopting business crept from 1.4 to 1.6. DSIT's survey of 3,500 businesses, conducted with IFF Research and Technopolis between February and May 2025, found that among adopters an average of 30 per cent of staff use AI, that 85 per cent of adopters use it for natural language and text generation, and, most tellingly, that most adopters report higher workforce productivity while most report no revenue change. Self-reported productivity with no revenue effect is the copilot problem in a single survey finding. If your organisation is drafting rules for these tools, our practical guide to AI governance policies for UK mid-market firms covers the policy side of the same problem.
What happens to the time a copilot saves
Freed time is redistributed by default, not banked by default, and the redistribution is invisible in company-wide averages. One contributor to the r/CIO thread, who should be read knowing they work for a workforce analytics vendor, put the measurement gap well: most organisations can say a copilot was deployed and is being used, but very few can say whether the hour saved on Tuesday became deeper work on Wednesday, more meetings on Thursday, or an earlier finish on Friday. Without that, the return conversation defaults to licence counts and sentiment surveys, and licence counts do not move a board.
The same contributor's sharper point survives the vendor discount: company-wide averages hide the signal. The identical tool produces different results on different teams. One team compounds the saved time into more shipped work; another absorbs it into coordination; a third works the same hours doing slightly less. Averaged together, these read as "no effect", which is precisely what the Danish administrative data shows at the aggregate level. Team-level measurement is where the answer lives, and team-level measurement only matters if a manager is accountable for deciding what the freed time becomes. Time that nobody reassigns, reassigns itself.
This is also why "AI will pay for itself in productivity" business cases keep getting harder to defend at renewal. The cost side is precise, per seat per month, visible on an invoice. The benefit side is a survey. Any board that has been through a few software renewals knows which side of that ledger to trust.
The board filter of pipeline, margin or risk
A board approves AI initiatives that move pipeline, margin or risk through a traceable metric, and rarely anything else. One r/CIO contributor, who runs an AI consultancy and was open about selling exactly this work, so weigh the framing accordingly, laid out the filter with worked examples that match what we see in UK deployments:
| Board lever | What qualifies | Example metrics |
|---|---|---|
| Revenue | Dynamic pricing, offer personalisation, automated lead routing | Conversion rate, basket size, lead acceptance rate |
| Cost | Automation that removes entire handoffs, not faster individual steps | Ticket triage through to resolution, claims adjudication, invoice matching |
| Risk | Compliance monitoring with audit trails, standardised vendor risk reviews | Findings surfaced, review cycle time, audit coverage |
The consultancy pitch does not make the filter wrong; finance directors have applied a version of it to every technology wave since client-server. What the filter usefully rules out is any initiative whose benefit statement is "people will be more productive". That claim, as the Danish evidence shows, is simultaneously true at the task level and unbankable at the company level.
The same contributor added a discipline worth stealing regardless of who sells it: without a before-and-after map of the organisation, time saved evaporates into busywork. Map the process owners, the service levels and the specific steps to be removed, then lock the gains into standard operating procedures and objectives. The mapping is not bureaucracy; it is the mechanism by which a saving becomes a number.
MIT's report points at the same misallocation from the spending side: more than half of generative AI budgets go to sales and marketing tools, while the research found the biggest returns in back-office automation, removing outsourced processing and external agency spend. The money is going where the demos are most impressive, not where the measurable savings are.
Fewer steps, not faster steps
Process-level AI clears boards because it removes handoffs rather than accelerating them, and a removed handoff is a saving you can point at. The phrase from the thread worth pinning above a whiteboard: "fewer steps, not faster steps". A copilot that helps a service desk agent write a ticket update faster is a faster step. A pipeline that takes ticket triage through to resolution without the ticket ever queueing for a human is a removed step, and removed steps have owners, service levels and costs that existed before and do not exist after. The before-and-after is auditable.
This maps onto what MIT found separates the successful 5 per cent: they pick one pain point, execute it well, and integrate deeply into a workflow rather than spreading a general-purpose tool across everyone. It also matches the purchasing pattern in the same research: buying from specialised vendors and partnering succeeded about 67 per cent of the time in MIT's dataset, while fully internal builds succeeded about a third as often. For a UK mid-market IT director the implication is uncomfortable but useful: the glamorous internal platform project is statistically the weakest option, and a narrow, bought, deeply integrated process tool is the strongest. The same logic, applied to suppliers rather than software, is why we argue UK MSPs should sell monthly outcomes rather than chase hardware margin: recurring, measurable service beats one-off acceleration.
One caution from the sceptics' side of the thread belongs here, because it is fair: a contributor described using AI to scale customer service during product recalls, capturing serial numbers, validating claims, reimbursing consumers, and another replied bluntly that this is ordinary automation and does not need AI at all. They are both right, and that is the point. If a process can be automated deterministically, automate it deterministically; the AI label adds cost and failure modes without adding capability. Which leads to the strongest idea in the whole discussion.
Use AI to build bridges, not to be the bridge
The most valuable framing in the r/CIO thread deserves to be quoted in spirit: use AI to build bridges, not to be the bridge. Putting AI permanently inside a process creates a permanent operating cost that rises with every transaction that crosses it, plus a permanent reliability question on every crossing. Using AI to design and build a deterministic system, one that was previously too complex or too time-consuming to specify, creates an asset. The bridge stands; the builder goes home.
The worked example from the thread makes it concrete. A network team had been stuck on a fault for two weeks. They gave an AI system read-only access and asked it to prove definitively what the issue was. It produced a diagnostic framework and the log queries to test it, and the whole exercise took under twenty minutes. The AI did not become part of the network's operating loop. It closed what that contributor called an effort gap rather than a skills gap: the valuable, tedious analysis nobody volunteers for. The output was understanding, and understanding, once produced, costs nothing to keep.
The same economics showed up in a smaller story elsewhere in the thread: one contributor built a Slack bot on a commercial AI API to let staff query CRM data conversationally, at a cost they put at a few hundred dollars a month for around 30 users, removing the need for a part-time hire. Note what made it work as a board story: a displaced cost with a name and a salary, not a productivity sentiment. These are practitioner accounts, not audited case studies, but the pattern they illustrate is the one the published research supports: the return appears where a specific cost or step disappears.
The bridge principle also gives you a clean test for any proposed agent: if this worked perfectly for six months, would we still need it running? If the answer is no, because the agent would by then have produced the runbook, the integration, the cleaned dataset or the deterministic script, then the agent is scaffolding and its cost should fall over time. If the answer is yes, you are buying a permanent per-transaction dependency, and the business case must carry that operating cost, plus oversight, forever. Some dependencies are worth it. Most are being bought without anyone noticing the distinction.
What material means when regulators are watching
Materiality is not only a finance judgement; for regulated and listed UK companies part of it is defined for you, and data sensitivity is where the definition bites first. One r/CIO contributor assesses AI materiality by the underlying data rather than the spend: searching public information about a competitor is one kind of decision; pointing a system at business-confidential data, personal data, health data or business-critical systems is a different kind entirely, whatever the licence cost.
UK law encodes a version of that instinct. Under Article 35 of the UK GDPR, a data protection impact assessment is mandatory before any processing that is likely to result in a high risk to individuals, and the ICO's criteria capture much of what AI deployments actually do: innovative technologies, systematic and extensive automated evaluation of people, large-scale use of special category data. The ICO's position is that if you are in doubt about whether the threshold is met, do the DPIA. In other words, for personal data there is a regulatory floor under "material", and an AI project can cross it at zero pounds of spend. The ICO notes this guidance is under review following the Data (Use and Access) Act, so check the current text before relying on it. Our guides to UK GDPR Article 30 records for cloud architects and how UK AI regulation compares with the EU AI Act cover the adjacent obligations, and the NCSC's actual guidance on ChatGPT and Copilot with sensitive business data is worth reading before your next security review, since it is widely misquoted in both directions.
For a listed company, the same logic runs through disclosure: a board weighing an AI initiative that touches business-critical systems or regulated data is weighing a risk item whether or not the spend clears any financial threshold. This is why the risk column of the board filter is underrated. Compliance monitoring with audit trails is unglamorous, but it speaks the board's native language.
The sceptics' case deserves a seat at the table
The sceptical replies in the thread are not noise; two of them describe failure modes that will consume real UK budgets this year. The first: much of what is being demonstrated as agentic AI on internal knowledge bases mainly demonstrates how disorganised the knowledge base was. An agent built on your standard operating procedures inherits their gaps, contradictions and staleness, and then presents them fluently. Organisations discovering this are, in effect, paying AI-era prices for a documentation audit they could have commissioned directly. That is not worthless, but it should be bought with open eyes, and often the durable asset is the cleaned documentation, not the agent. The bridge principle again.
The second: some AI adoption is investor-facing rather than value-creating. The same contributor observed that in 2026 having AI features on the website and in the pitch deck carries value in itself, because investors are wary of companies with no AI story, regardless of operational impact. A CIO should at least be honest internally about which category a given initiative belongs to. Signalling spend is a legitimate board decision; signalling spend mislabelled as a productivity programme is how the 95 per cent gets populated.
On reliability, one thread contributor argued that generative systems are strong at retrieving from a provided dataset and degrade sharply when asked to reason beyond it, and offered hallucination-rate figures of roughly 1.7 to 3 per cent for the best systems and up to 20 per cent for the worst. Those figures are one practitioner's claim, unsourced in the thread, and we could not tie them to a citable benchmark, so treat the numbers as anecdote. The design instinct behind them, however, is sound and matches the published RCT literature's warning that effects turn negative when the tools are applied to the wrong tasks: constrain generative components to retrieval and drafting inside a checked process, and keep anything requiring guaranteed accuracy deterministic. Teams that want the retrieval benefits without a per-seat SaaS dependency increasingly look at self-hosted alternatives, which at least make the operating cost explicit.
MIT's research adds one more sceptical data point worth carrying into procurement: the widespread "shadow AI" economy, where staff use unsanctioned consumer tools that outperform the official deployment. If your sanctioned copilot is losing to a personal ChatGPT subscription, the licence report your board sees is measuring the wrong thing twice. Controlling what these tools can reach, as we set out in stopping Microsoft 365 Copilot surfacing confidential data to the wrong people, matters more than the adoption dashboard.
How to build an AI case a board will pass
Build the case backwards from the metric, and be candid about which of the three levers it moves. A sequence that follows from everything above:
- Name the lever. Pipeline, margin or risk, with the specific metric the board already sees. If no metric fits, the initiative may still be worth doing, but present it as capability-building or signalling, not return.
- Map before and after. Owners, service levels, and the steps that will no longer exist. "Fewer steps, not faster steps" is the test; a case with no removed steps is a productivity case, and productivity cases need the next item to survive.
- Pre-assign the freed time. If the benefit is time saved, name the team, the manager accountable, and what the hours become. Measure at team level, because averages will hide whatever happens. Unassigned time is not a benefit; it is a hope.
- Decide bridge or bridge-builder. Will this system still need to run once it has worked? If not, plan for its cost to fall to zero and its output to become a deterministic asset. If yes, carry the rising operating and oversight cost in the case honestly.
- Buy narrow before building broad. MIT's 67 per cent success rate for bought, specialised, deeply integrated tools against internal builds succeeding a third as often is the strongest procurement statistic in this field. Treat an internal platform build as the exceptional choice needing exceptional justification.
- Clear the regulatory floor first. Personal data, special category data or business-critical systems put you in DPIA territory under Article 35 UK GDPR before any financial threshold is reached. Do that assessment before the business case, not after the pilot.
- Sunset by default. Give every pilot a date on which it must show its metric moved or stop. The 95 per cent in MIT's data are not mostly cancelled projects; they are pilots that never ended.
None of this is hostile to copilots. A 3 per cent time saving across a workforce is real value if, and only if, someone decides what the 3 per cent becomes. The organisations in the successful 5 per cent are not the ones with better models. They are the ones where the freed time and the removed steps have names against them.
Sources
This article draws on a June 2026 practitioner discussion among technology leaders on r/CIO (20 points, 19 comments), whose accounts are attributed throughout as individual reports, with two contributors' vendor affiliations disclosed where their points are used. Published evidence comes from MIT NANDA's report The GenAI Divide, State of AI in Business 2025 (August 2025, based on 150 leader interviews, a 350-employee survey and 300 public deployments); the working paper Large Language Models, Small Labor Market Effects by Anders Humlum (University of Chicago Booth) and Emilie Vestergaard (University of Copenhagen), July 2025; the Office for National Statistics bulletin Artificial Intelligence in UK Businesses 2023 to 2026, released 20 July 2026; DSIT's AI Adoption Research conducted by IFF Research and Technopolis Group, updated 13 February 2026; and the ICO's guidance on when a data protection impact assessment is required under Article 35 UK GDPR.