I have lost count of how many times in the last few months a client, a former colleague or someone I have not spoken to in years has messaged me with some version of the same question. How do we actually measure ROI on our AI spend. Not should we do AI, everyone has moved past that.
The real question now is sharper & quite uncomfortable.
We have spent the money, the model works in testing, so why can nobody tell me what we got back for it.
Every time, my answer starts in the same place and it is rarely where they expect. The ROI problem is almost never a measurement problem.
It is a spending problem that happened months earlier, before anyone thought to ask what to measure. There is a pattern so common in AI programs today that you could set your watch by it. A department or company announces an AI initiative.
A model gets built or bought.
A demo dazzles the board or the minister.
Six months later, the model still works fine. Almost nobody is using it. And now someone wants a number to put in a slide.
This is not a technology failure. It is a budgeting failure and a research finding from Boston Consulting Group explains why, in numbers simple enough to fit on a napkin.
The rule, in one sentence
BCG’s research suggests that successful AI transformation depends on dedicating roughly 10% of the effort to algorithms, 20% to technology and data, and the remaining 70% to people and process. Not the other way around, which is how most budgets are actually built.
Put plainly: the model is the tip of the iceberg. The part that decides whether the initiative lives or dies, training, workflow redesign, incentives, trust, the willingness of a claims officer or a caseworker to actually change how they work, sits underwater, unglamorous and chronically underfunded. If your budget spent 80% of its money on the 10%, do not be surprised when the ROI calculation comes back looking thin. You measured the wrong investment.
This isn’t an opinion. It’s a pattern that keeps repeating. Research shows only a small fraction of managers currently have strong AI skills, and roughly one in four employees shows real fluency with generative AI tools, with most reporting their training has been inadequate. The tool is rarely the bottleneck. The organization around the tool is.
That framework was built for the first wave of enterprise AI: recommendation engines, forecasting models, fraud detection. It is even more true now, and it needs a rewrite for the world we are actually in. One of large language models, AI agents that take actions rather than just predictions, and, especially in the public sector, a citizen watching over your shoulder.
Why 2026 changes the maths, not the message
Three things have shifted since the original framework was written.
The algorithm isn’t one thing anymore.
In 2023, an AI project meant training a model on your data. In 2026, it more often means renting a foundation model built by someone else and wrapping it around your systems and your documents. The 10% has gotten cheaper and faster to stand up. That’s good news except for one problem: organizations now blow through the easy part in weeks and arrive at the hard 70% having budgeted no time or money for it at all.
Tech and data now includes something the original 20% never had to reckon with: Agents that act rather than just answer.
A chatbot that drafts a memo is low stakes. A system that decides which files get flagged for audit, which applicants get fast tracked or which citizen call gets escalated is a different category of infrastructure risk. The 20% today has to include monitoring, audit trails and kill switches, not just pipelines and storage.
This is the one enterprises can mostly skip and governments cannot, there is a fourth force sitting on top of the original three: legitimacy.
A private company that ships a flawed recommendation engine loses some revenue. A government agency that ships a flawed AI decision on benefits, policing or immigration loses something much harder to rebuild: public consent to be governed by a machine at all. The old framework had no line item for that. The new one needs one.
The adapted framework: 10-20-70, plus the number nobody budgets for
Keep the same skeleton. Redefine what lives inside it. Add the piece the original framework didn’t need to price in.
10%, the model.
Choosing, fine tuning or prompting the system.
This is now the fastest and cheapest part of any AI program, which is precisely why it eats a disproportionate share of the headlines and the budget. Treat it as a commodity decision, not the centerpiece of your strategy.
20%, the plumbing.
Data quality, integration with legacy systems, security and, new to this decade, the guardrails that let an AI agent take an action safely: logging, human checkpoints and the ability to pause or reverse a decision. Governments increasingly recognize that data readiness is the critical precondition for their AI ambitions, and it shows: this is usually where public sector projects quietly die, buried under decades of systems that were never built to talk to each other.
70%, the human machine.
Retraining the workforce, redesigning the actual workflow instead of bolting AI onto the old one, rewriting incentives so people are rewarded for using the tool well rather than punished for the mistakes it makes, and building the muscle of managers who can coach a team through the change. Real impact requires more than isolated pilots. It requires redesigning processes, ways of working and operating models so the technology becomes part of everyday service delivery, not a bolt on.
The plus one, trust.
Not a percentage so much as a precondition that multiplies or zeroes out everything above it. Explainability, accountability for when the system gets it wrong, and transparency with the public about where AI is and isn’t making the call. Skip this layer in a bank and you get complaints. Skip it in a government agency and you get a headline, a parliamentary inquiry and a program that gets cancelled regardless of how good the model was.
The evidence this isn’t theoretical
Look at what’s actually happening across public sector AI programs right now and the framework stops being an abstraction.
Only about 28% of public sector organizations have successfully scaled their AI initiatives beyond the pilot stage. That is not a technology gap. Most of these pilots work. It’s a 70% gap.
Nearly half of governments plan to deploy AI at scale within the next year, yet more than four in ten admit they are already hitting roadblocks scaling their current initiatives. And tellingly, these same organizations report that around 40% of their technology budgets still go toward simply maintaining existing systems, meaning the 20% is being starved even before anyone gets to the 70%.
Even the diagnosis of why pilots stall keeps landing in the same place. Governments’ AI initiatives commonly fail to move past the pilot stage because of problems with workflow integration, data access and operating costs, not because the underlying models don’t work.
None of this is a public sector problem specifically. It is what happens anywhere an organization buys the 10% and skips the 70%, and then wonders why the ROI math will not close.
What this means if you’re the one writing the cheque
For a leader, whether minister, agency head, CEO or chief data officer, the adapted framework is really a budgeting discipline disguised as a percentage, and it changes what you should be asking your team to measure.
Stop measuring AI programs by how many models you’ve deployed.
Measure them by how many workflows have actually changed and how many frontline staff can tell you, unprompted, how the tool changed their Tuesday. A model nobody’s workflow depends on is a very expensive science project, and no amount of clever ROI formula will make it otherwise.
Fund the boring 70% before you fund the exciting 10%.
In practice this means budgeting for training and change management from day one, not as an afterthought once the model ships. Put a manager, not just a data scientist, in charge of the rollout. The evidence backs this instinct directly: peer learning from colleagues, not formal training programs, is consistently the primary way people actually acquire AI skills, which tells you the investment that matters most is time for people to learn from each other, not another license.
Treat trust as an engineering requirement, not a communications exercise.
If a citizen or a customer cannot get a clear answer to why the system decided something about them, the program is not finished, no matter how accurate the model is.
Sequence deliberately
McKinsey’s own read on government AI programs lands on a similar four step logic: start from the mission outcome you actually want rather than the technology, redesign the end to end workflow around it, build an operating model suited to AI and keep a human in the loop for consequential decisions. That order matters. Reverse it and start with the technology, and you get exactly the stalled pilots showing up in every survey this year, and exactly the ROI conversations that go nowhere.
What this means if you’re the one building it
For the practitioners, the data scientists, engineers and program leads actually shipping these systems, the framework is a warning against your own instincts. The interesting work is in the 10%. The tempting metric is model accuracy. The comfortable conversation is about architecture.
But the job that actually determines whether your work survives contact with a real organization is the boring 70%: sitting with the caseworker who will use this tool and watching where it breaks their existing habits, building the dashboard a manager will actually check, writing the one page explanation a citizen can read without a data science degree.
The single best predictor of whether an AI program survives its first year isn’t the benchmark score of the model behind it. It’s whether anyone spent real time and money on the 70% that has nothing to do with the model at all.
So the next time someone asks me how to evaluate ROI on their AI program, my first question back is not about the model at all. It is where their budget actually went.
That was true when BCG first wrote it down.
It is more true now, with agents that act rather than just advise, and a public that is watching more closely than ever whether the machines making decisions about their lives can be trusted, understood and, when necessary, overruled.


