Which AI Projects Actually Clear the Board, and Why Copilots Keep Falling Short

Which AI Projects Actually Clear the Board, and Why Copilots Keep Falling Short

17 min read

Most AI investment goes into individual-productivity tools, yet boards approve initiatives that move pipeline, margin or risk through a metric they already track. Drawing on a June 2026 discussion among technology leaders and published evidence from MIT NANDA, Danish administrative data, the ONS and DSIT, this piece shows why time saved by copilots is redistributed rather than banked, why company-wide averages hide the signal, and why process-level AI that removes steps clears approval. It sets out the bridge-builder principle, the regulatory floor under materiality for UK companies, and a seven-step method for building an AI case a board will pass.

Daniel Thomas
Written by Daniel Thomas

The AI projects that clear a board are the ones that touch pipeline, margin or risk through a metric the board already tracks. The ones that stall are overwhelmingly individual-productivity deployments, copilots and assistants bought seat by seat, because nobody can say what happened to the time they saved. That is not a cynical reading; it is what the published evidence now shows. MIT's NANDA initiative reported in August 2025 that about 95 per cent of enterprise generative AI pilots deliver little or no measurable profit-and-loss impact, while roughly 5 per cent achieve rapid revenue acceleration. The Department for Science, Innovation and Technology found in research updated in February 2026 that most UK businesses using AI report a productivity increase, yet most have seen no change in revenue. Productivity that never becomes revenue, cost or reduced risk is exactly the gap this article is about, and it is a gap with a cause that can be fixed.

The argument in brief, drawn from a June 2026 discussion among technology leaders on r/CIO and tested against published research below: individual-productivity AI accelerates existing work, and accelerated work is invisible to a board unless someone deliberately reassigns the time it frees. Process-level AI redesigns the work itself, which forces the reassignment to happen and makes the return legible. The distinction sounds subtle. Financially, it is the whole game.

Why individual-productivity AI keeps failing the board test

Individual-productivity AI fails the board test because organisations measure deployment and usage, not outcomes. A CIO can report that 4,000 copilot licences are active and adoption is at 70 per cent, and none of that is a business result. The question a finance director will ask, and the one almost nobody can answer, is the one posed in that r/CIO thread: suppose the copilot works, and a task that took a week now takes a day. What happened to the other four days? Were they reassigned to something more valuable, absorbed into meetings and low-impact work, or nothing at all? If the organisation cannot answer, the benefit cannot be booked.

The best field evidence available says the honest answer is usually "not much". Economists Anders Humlum and Emilie Vestergaard linked two large adoption surveys, covering roughly 25,000 workers across 7,000 workplaces in 11 exposed occupations, to Danish administrative records on earnings and hours. Their July 2025 working paper found average time savings from AI chatbots of about 3 per cent of working hours, and precisely estimated null effects on earnings and hours, with confidence intervals ruling out effects above 1 per cent. This in occupations, from software development to legal work, where randomised controlled trials had shown productivity gains of 15 to 50 per cent. The tools work in the lab. In the wild, the saved time dissolves, partly into new tasks the tools themselves create, which the authors describe as integration and oversight work.

The Danish study also found that 38 per cent of the surveyed workplaces had deployed enterprise chatbot versions and 30 per cent of employees had received training, so this is not a story about immature adoption. Employers invested; the investment raised usage; usage did not move the numbers a board reads.

UK survey data rhymes with this. The Office for National Statistics reported on 20 July 2026 that self-reported AI use among UK businesses with ten or more employees rose from around 12 per cent to around 35 per cent since late 2023, but describes adoption as shallow: the average number of AI technologies per adopting business crept from 1.4 to 1.6. DSIT's survey of 3,500 businesses, conducted with IFF Research and Technopolis between February and May 2025, found that among adopters an average of 30 per cent of staff use AI, that 85 per cent of adopters use it for natural language and text generation, and, most tellingly, that most adopters report higher workforce productivity while most report no revenue change. Self-reported productivity with no revenue effect is the copilot problem in a single survey finding. If your organisation is drafting rules for these tools, our practical guide to AI governance policies for UK mid-market firms covers the policy side of the same problem.

What happens to the time a copilot saves

Freed time is redistributed by default, not banked by default, and the redistribution is invisible in company-wide averages. One contributor to the r/CIO thread, who should be read knowing they work for a workforce analytics vendor, put the measurement gap well: most organisations can say a copilot was deployed and is being used, but very few can say whether the hour saved on Tuesday became deeper work on Wednesday, more meetings on Thursday, or an earlier finish on Friday. Without that, the return conversation defaults to licence counts and sentiment surveys, and licence counts do not move a board.

The same contributor's sharper point survives the vendor discount: company-wide averages hide the signal. The identical tool produces different results on different teams. One team compounds the saved time into more shipped work; another absorbs it into coordination; a third works the same hours doing slightly less. Averaged together, these read as "no effect", which is precisely what the Danish administrative data shows at the aggregate level. Team-level measurement is where the answer lives, and team-level measurement only matters if a manager is accountable for deciding what the freed time becomes. Time that nobody reassigns, reassigns itself.

This is also why "AI will pay for itself in productivity" business cases keep getting harder to defend at renewal. The cost side is precise, per seat per month, visible on an invoice. The benefit side is a survey. Any board that has been through a few software renewals knows which side of that ledger to trust.

The board filter of pipeline, margin or risk

A board approves AI initiatives that move pipeline, margin or risk through a traceable metric, and rarely anything else. One r/CIO contributor, who runs an AI consultancy and was open about selling exactly this work, so weigh the framing accordingly, laid out the filter with worked examples that match what we see in UK deployments:

Board leverWhat qualifiesExample metrics
RevenueDynamic pricing, offer personalisation, automated lead routingConversion rate, basket size, lead acceptance rate
CostAutomation that removes entire handoffs, not faster individual stepsTicket triage through to resolution, claims adjudication, invoice matching
RiskCompliance monitoring with audit trails, standardised vendor risk reviewsFindings surfaced, review cycle time, audit coverage

The consultancy pitch does not make the filter wrong; finance directors have applied a version of it to every technology wave since client-server. What the filter usefully rules out is any initiative whose benefit statement is "people will be more productive". That claim, as the Danish evidence shows, is simultaneously true at the task level and unbankable at the company level.

The same contributor added a discipline worth stealing regardless of who sells it: without a before-and-after map of the organisation, time saved evaporates into busywork. Map the process owners, the service levels and the specific steps to be removed, then lock the gains into standard operating procedures and objectives. The mapping is not bureaucracy; it is the mechanism by which a saving becomes a number.

MIT's report points at the same misallocation from the spending side: more than half of generative AI budgets go to sales and marketing tools, while the research found the biggest returns in back-office automation, removing outsourced processing and external agency spend. The money is going where the demos are most impressive, not where the measurable savings are.

Fewer steps, not faster steps

Process-level AI clears boards because it removes handoffs rather than accelerating them, and a removed handoff is a saving you can point at. The phrase from the thread worth pinning above a whiteboard: "fewer steps, not faster steps". A copilot that helps a service desk agent write a ticket update faster is a faster step. A pipeline that takes ticket triage through to resolution without the ticket ever queueing for a human is a removed step, and removed steps have owners, service levels and costs that existed before and do not exist after. The before-and-after is auditable.

This maps onto what MIT found separates the successful 5 per cent: they pick one pain point, execute it well, and integrate deeply into a workflow rather than spreading a general-purpose tool across everyone. It also matches the purchasing pattern in the same research: buying from specialised vendors and partnering succeeded about 67 per cent of the time in MIT's dataset, while fully internal builds succeeded about a third as often. For a UK mid-market IT director the implication is uncomfortable but useful: the glamorous internal platform project is statistically the weakest option, and a narrow, bought, deeply integrated process tool is the strongest. The same logic, applied to suppliers rather than software, is why we argue UK MSPs should sell monthly outcomes rather than chase hardware margin: recurring, measurable service beats one-off acceleration.

One caution from the sceptics' side of the thread belongs here, because it is fair: a contributor described using AI to scale customer service during product recalls, capturing serial numbers, validating claims, reimbursing consumers, and another replied bluntly that this is ordinary automation and does not need AI at all. They are both right, and that is the point. If a process can be automated deterministically, automate it deterministically; the AI label adds cost and failure modes without adding capability. Which leads to the strongest idea in the whole discussion.

Use AI to build bridges, not to be the bridge

The most valuable framing in the r/CIO thread deserves to be quoted in spirit: use AI to build bridges, not to be the bridge. Putting AI permanently inside a process creates a permanent operating cost that rises with every transaction that crosses it, plus a permanent reliability question on every crossing. Using AI to design and build a deterministic system, one that was previously too complex or too time-consuming to specify, creates an asset. The bridge stands; the builder goes home.

The worked example from the thread makes it concrete. A network team had been stuck on a fault for two weeks. They gave an AI system read-only access and asked it to prove definitively what the issue was. It produced a diagnostic framework and the log queries to test it, and the whole exercise took under twenty minutes. The AI did not become part of the network's operating loop. It closed what that contributor called an effort gap rather than a skills gap: the valuable, tedious analysis nobody volunteers for. The output was understanding, and understanding, once produced, costs nothing to keep.

The same economics showed up in a smaller story elsewhere in the thread: one contributor built a Slack bot on a commercial AI API to let staff query CRM data conversationally, at a cost they put at a few hundred dollars a month for around 30 users, removing the need for a part-time hire. Note what made it work as a board story: a displaced cost with a name and a salary, not a productivity sentiment. These are practitioner accounts, not audited case studies, but the pattern they illustrate is the one the published research supports: the return appears where a specific cost or step disappears.

The bridge principle also gives you a clean test for any proposed agent: if this worked perfectly for six months, would we still need it running? If the answer is no, because the agent would by then have produced the runbook, the integration, the cleaned dataset or the deterministic script, then the agent is scaffolding and its cost should fall over time. If the answer is yes, you are buying a permanent per-transaction dependency, and the business case must carry that operating cost, plus oversight, forever. Some dependencies are worth it. Most are being bought without anyone noticing the distinction.

What material means when regulators are watching

Materiality is not only a finance judgement; for regulated and listed UK companies part of it is defined for you, and data sensitivity is where the definition bites first. One r/CIO contributor assesses AI materiality by the underlying data rather than the spend: searching public information about a competitor is one kind of decision; pointing a system at business-confidential data, personal data, health data or business-critical systems is a different kind entirely, whatever the licence cost.

UK law encodes a version of that instinct. Under Article 35 of the UK GDPR, a data protection impact assessment is mandatory before any processing that is likely to result in a high risk to individuals, and the ICO's criteria capture much of what AI deployments actually do: innovative technologies, systematic and extensive automated evaluation of people, large-scale use of special category data. The ICO's position is that if you are in doubt about whether the threshold is met, do the DPIA. In other words, for personal data there is a regulatory floor under "material", and an AI project can cross it at zero pounds of spend. The ICO notes this guidance is under review following the Data (Use and Access) Act, so check the current text before relying on it. Our guides to UK GDPR Article 30 records for cloud architects and how UK AI regulation compares with the EU AI Act cover the adjacent obligations, and the NCSC's actual guidance on ChatGPT and Copilot with sensitive business data is worth reading before your next security review, since it is widely misquoted in both directions.

For a listed company, the same logic runs through disclosure: a board weighing an AI initiative that touches business-critical systems or regulated data is weighing a risk item whether or not the spend clears any financial threshold. This is why the risk column of the board filter is underrated. Compliance monitoring with audit trails is unglamorous, but it speaks the board's native language.

The sceptics' case deserves a seat at the table

The sceptical replies in the thread are not noise; two of them describe failure modes that will consume real UK budgets this year. The first: much of what is being demonstrated as agentic AI on internal knowledge bases mainly demonstrates how disorganised the knowledge base was. An agent built on your standard operating procedures inherits their gaps, contradictions and staleness, and then presents them fluently. Organisations discovering this are, in effect, paying AI-era prices for a documentation audit they could have commissioned directly. That is not worthless, but it should be bought with open eyes, and often the durable asset is the cleaned documentation, not the agent. The bridge principle again.

The second: some AI adoption is investor-facing rather than value-creating. The same contributor observed that in 2026 having AI features on the website and in the pitch deck carries value in itself, because investors are wary of companies with no AI story, regardless of operational impact. A CIO should at least be honest internally about which category a given initiative belongs to. Signalling spend is a legitimate board decision; signalling spend mislabelled as a productivity programme is how the 95 per cent gets populated.

On reliability, one thread contributor argued that generative systems are strong at retrieving from a provided dataset and degrade sharply when asked to reason beyond it, and offered hallucination-rate figures of roughly 1.7 to 3 per cent for the best systems and up to 20 per cent for the worst. Those figures are one practitioner's claim, unsourced in the thread, and we could not tie them to a citable benchmark, so treat the numbers as anecdote. The design instinct behind them, however, is sound and matches the published RCT literature's warning that effects turn negative when the tools are applied to the wrong tasks: constrain generative components to retrieval and drafting inside a checked process, and keep anything requiring guaranteed accuracy deterministic. Teams that want the retrieval benefits without a per-seat SaaS dependency increasingly look at self-hosted alternatives, which at least make the operating cost explicit.

MIT's research adds one more sceptical data point worth carrying into procurement: the widespread "shadow AI" economy, where staff use unsanctioned consumer tools that outperform the official deployment. If your sanctioned copilot is losing to a personal ChatGPT subscription, the licence report your board sees is measuring the wrong thing twice. Controlling what these tools can reach, as we set out in stopping Microsoft 365 Copilot surfacing confidential data to the wrong people, matters more than the adoption dashboard.

How to build an AI case a board will pass

Build the case backwards from the metric, and be candid about which of the three levers it moves. A sequence that follows from everything above:

  1. Name the lever. Pipeline, margin or risk, with the specific metric the board already sees. If no metric fits, the initiative may still be worth doing, but present it as capability-building or signalling, not return.
  2. Map before and after. Owners, service levels, and the steps that will no longer exist. "Fewer steps, not faster steps" is the test; a case with no removed steps is a productivity case, and productivity cases need the next item to survive.
  3. Pre-assign the freed time. If the benefit is time saved, name the team, the manager accountable, and what the hours become. Measure at team level, because averages will hide whatever happens. Unassigned time is not a benefit; it is a hope.
  4. Decide bridge or bridge-builder. Will this system still need to run once it has worked? If not, plan for its cost to fall to zero and its output to become a deterministic asset. If yes, carry the rising operating and oversight cost in the case honestly.
  5. Buy narrow before building broad. MIT's 67 per cent success rate for bought, specialised, deeply integrated tools against internal builds succeeding a third as often is the strongest procurement statistic in this field. Treat an internal platform build as the exceptional choice needing exceptional justification.
  6. Clear the regulatory floor first. Personal data, special category data or business-critical systems put you in DPIA territory under Article 35 UK GDPR before any financial threshold is reached. Do that assessment before the business case, not after the pilot.
  7. Sunset by default. Give every pilot a date on which it must show its metric moved or stop. The 95 per cent in MIT's data are not mostly cancelled projects; they are pilots that never ended.

None of this is hostile to copilots. A 3 per cent time saving across a workforce is real value if, and only if, someone decides what the 3 per cent becomes. The organisations in the successful 5 per cent are not the ones with better models. They are the ones where the freed time and the removed steps have names against them.

Sources

This article draws on a June 2026 practitioner discussion among technology leaders on r/CIO (20 points, 19 comments), whose accounts are attributed throughout as individual reports, with two contributors' vendor affiliations disclosed where their points are used. Published evidence comes from MIT NANDA's report The GenAI Divide, State of AI in Business 2025 (August 2025, based on 150 leader interviews, a 350-employee survey and 300 public deployments); the working paper Large Language Models, Small Labor Market Effects by Anders Humlum (University of Chicago Booth) and Emilie Vestergaard (University of Copenhagen), July 2025; the Office for National Statistics bulletin Artificial Intelligence in UK Businesses 2023 to 2026, released 20 July 2026; DSIT's AI Adoption Research conducted by IFF Research and Technopolis Group, updated 13 February 2026; and the ICO's guidance on when a data protection impact assessment is required under Article 35 UK GDPR.

Data & Insights

What MIT found enterprise generative AI pilots actually deliver

MIT's NANDA initiative reported that about 95 per cent of enterprise generative AI pilots deliver little or no measurable profit-and-loss impact, while roughly 5 per cent achieve rapid revenue acceleration.

Source: MIT NANDA initiative, August 2025

Frequently Asked Questions

Why do most AI pilots fail to show a return?

Most AI pilots fail to show a return because they accelerate individual work without anyone deciding what the freed time becomes, so the benefit never reaches a metric the board tracks. MIT NANDA's August 2025 research found about 95 per cent of enterprise generative AI pilots deliver little or no measurable profit-and-loss impact, and attributes the failures to weak workflow integration rather than model quality. Danish field data tells the same story from the other end: about 3 per cent of working time saved, with no measurable effect on hours or earnings.

What AI projects do boards actually approve?

Boards approve AI projects that move pipeline, margin or risk through a traceable metric they already see. In practice that means revenue work such as dynamic pricing and lead routing, cost work that removes entire handoffs such as ticket triage through to resolution or invoice matching, and risk work such as compliance monitoring with audit trails. Initiatives whose benefit statement is only that people will be more productive rarely qualify, because the productivity cannot be traced to a booked number.

What happens to the time a copilot saves?

Time a copilot saves is redistributed by default rather than banked. It becomes deeper work on some teams, more meetings on others, and slightly shorter effective hours elsewhere, and company-wide averages blur these into no visible effect. The Humlum and Vestergaard study of 25,000 Danish workers found average time savings of about 3 per cent and null effects on earnings and hours. The saved time only becomes value when a named manager reassigns it and the result is measured at team level.

What does material mean for an AI project in a UK company?

Material has both a financial and a regulatory meaning for UK companies. Financially, an initiative is material when it moves pipeline, margin or risk in a way the board can trace. Under Article 35 of the UK GDPR, materiality also has a floor set by data sensitivity: processing likely to result in high risk to individuals, which covers much AI work on personal or special category data, requires a data protection impact assessment before it starts, regardless of how small the spend is.

Is it better to buy AI tools or build them internally?

The published evidence favours buying narrow, specialised tools over building internally. MIT NANDA found purchased tools from specialised vendors succeeded about 67 per cent of the time, while internal builds succeeded about a third as often. The successful minority of adopters pick one pain point, buy or partner for a tool that integrates deeply into that workflow, and give line managers ownership of adoption, rather than spreading a general-purpose assistant across the whole organisation.

What does use AI to build bridges, not to be the bridge mean?

It means using AI to create deterministic systems rather than placing AI permanently inside a process. An AI kept in the loop is a bridge: every transaction crosses it, so its cost and its reliability risk rise with use, forever. AI used as a bridge-builder produces an asset, such as a diagnosis, a runbook, a cleaned dataset or a script, and then steps away, so the cost falls to zero once the work is done. A useful test is whether the system would still be needed after six months of working perfectly.

How should CIOs measure AI productivity gains?

CIOs should measure AI productivity at team level, against a pre-assigned use for the freed time, not through licence counts or sentiment surveys. Deployment and usage statistics show activity, not value. The workable pattern is to name the team and accountable manager, state in advance what the saved hours become, and compare team output before and after. Company-wide averages hide the signal because different teams absorb freed time differently, which is why aggregate studies keep finding near-zero net effects.

How many UK businesses are actually using AI?

Around 35 per cent of UK businesses with ten or more employees reported using AI by mid 2026, up from around 12 per cent in late 2023, according to the ONS bulletin of 20 July 2026. Adoption is shallow, averaging 1.6 AI technologies per adopting business, and uneven, from 58 per cent in information and communication to 13 per cent in construction. DSIT's separate 2025 survey, using a tighter definition, put adoption at one in six businesses, with 30 per cent of staff using AI on average within adopters.

Do AI copilots increase revenue?

There is no good evidence yet that individual copilots increase revenue. DSIT's research updated in February 2026 found most UK businesses using AI report a productivity increase while most report no change in revenue. Danish administrative data shows the same gap: real task-level time savings that never appear in earnings or hours. Revenue effects show up where AI is applied to revenue processes directly, such as pricing, personalisation and lead routing, with the metric defined before deployment.

When does an AI project need a DPIA in the UK?

An AI project needs a data protection impact assessment when its processing of personal data is likely to result in a high risk to individuals, under Article 35 UK GDPR. The ICO's triggers include use of innovative technologies, systematic and extensive automated evaluation of people, and large-scale use of special category data, which between them cover much enterprise AI. The ICO advises doing a DPIA whenever in doubt, and notes its guidance is under review following the Data (Use and Access) Act.

Are agents built on internal documentation worth it?

Agents built on internal documentation are worth it mainly when the documentation is sound, and one technology leader's observation cuts to the risk: impressive-looking agents often mostly reveal how disorganised the underlying knowledge base was. The agent inherits gaps, contradictions and stale content, then presents them fluently. Often the durable asset from such a project is the cleaned documentation rather than the agent itself, so it can be rational to treat the exercise as a documentation overhaul with an AI interface as a by-product.

Why do companies keep buying AI that does not pay back?

Partly because measurement stops at deployment, and partly because some AI adoption is aimed at investors rather than operations. One contributor to the June 2026 discussion observed that in 2026 visible AI capability carries value with investors regardless of operational impact. That can be a legitimate signalling decision, but when signalling spend is presented internally as a productivity programme it populates the 95 per cent of pilots that never show a measurable return, and it erodes the credibility of the cases that would.