You've seen the headlines. You've overheard the water-cooler version too: AI is coming for the jobs. And in plenty of cases it can. There are real redundancies in how a lot of work gets done across industries. But here's my problem with the conversation. Organizations, the media, all of it, are over-indexing on what AI can do and will do, and most of them can't actually measure the true cost of what they're implementing, because they were never measuring the outputs you'd need to justify that cost in the first place. They were simply running with the narrative with no real strategy behind it. Worse yet, the individuals I've seen implementing AI are not experts in AI nor do they have experience in organizational transformation and process improvement. Dare I say most organizations and individuals are reacting to the AI frenzy just to win favor with investors and customers just so they can say they are innovative.

So let me show you what happens when someone actually puts a number on it. The clearest place to watch this play out is in software development, because coding is one of the few jobs where you can measure the input and output. So people did.

In 2023, GitHub published a study that made the case for AI at work about as cleanly as anyone has. Developers were asked to build one thing, a web server, from scratch. Half had GitHub's AI assistant, half didn't. The group with AI finished 55% faster. Not a rounding error. More than twice the speed, measured with a stopwatch.

That's the number that put AI in everyone's budget. Worth noticing who published it: GitHub, the company that sells the assistant. It's a real result, and it's also a sales figure. Hold onto both of those.

Now the study nobody put on a slide.

In 2025, an independent group called METR asked the same basic question, but built the test to be realistic instead of clean. Not a web server from scratch, but sixteen seasoned developers working 246 real tasks inside their own mature codebases, the kind of code they've lived in for years. For each task, a coin flip decided whether they could use AI, and METR timed all of it. Different study, different developers, harder conditions, one shared question: did AI actually make them faster?

Every one of them said yes. They put it around 20% faster, and they were sure of it.

The actual result was 19% slower.

The gap between those two studies is the whole story, and it isn't that the tool got worse between 2023 and 2025. It's who was using it, and on what. On a clean, self-contained task, especially in less-experienced hands, AI genuinely speeds you up, which is exactly what GitHub measured. Drop that same tool into a senior engineer's mature codebase, and the prompting, the waiting, and the cleanup to make its suggestions actually fit can cost more time than they save.

And that's the over-indexing problem made concrete. It isn't that AI doesn't pay off. It's that the people using it often can't even tell whether it did. If a senior engineer can be that wrong about their own work, in the wrong direction, the productivity number they hand up the chain, the one that justifies the spend, was simply unreliable.

What the numbers actually say

That was coding. One industry, one kind of task. Step back and look at the whole field and the same thing happens everywhere: the return depends completely on who's using AI and what they're using it on. Which is a problem, because almost everyone out there is trying to tell you how much time AI can save you. Listen, AI has saved me a lot of time in the areas which are important to me. However, using AI can only be as strong as the person wielding it.

Start with a saving figure you've either heard or probably had quoted to you. A Microsoft-commissioned study puts the average return on AI at 3.5x what you put in, and as high as 8x. And my problem with it isn't just that the company selling the software also ran the math, though that should bug you. It's that word, average. It takes a return that swings all over the place and flattens it into one confident multiple, like AI pays off the same for everybody. It won't, and the independent studies show you exactly how much it doesn't.

Economists at NBER, an independent research group, tracked 5,179 customer-support agents using an AI assistant and measured one concrete thing: issues resolved per hour. The newest agents got 34% faster, because the AI handed them the knowledge the veterans had spent years earning. The veterans themselves? Almost nothing. Same tool, same floor, same day, and the return ran from huge to zero based on who sat down at the desk.

So, here's the honest range. Novice on a simple, countable task, big win. Expert on messy, high-context work, break-even or a loss. That is not a rounding error around 3.5x, that is the return being conditional, top to bottom. Which means the question everyone keeps asking, what's the ROI on AI, is the wrong question. The question that should be asked is how AI will improve efficiency across various workflows which are unique to my business.

How people are actually using AI

If it all comes down to who's using AI and on what, it's worth stepping back from the studies to see how people are actually using it day to day.

When Pew asked Americans what they use AI for, the answers were very interesting but aligned with my general thoughts. Search, 42%. Work tasks, 38%. Entertainment, 25%. Making images, 24%. Medical advice, 20%. Emotional support, 10%. People are using AI to look things up, generate funny images to send to your grandma and generally to get through the day. A good chunk of it never touches a business process at all.

That emotional-support number is really a fascinating study. People leaning on AI for the questions they'd never ask another person. Honestly, I've seen people create AI boyfriends and girlfriends. This is an entire story on its own, and honestly an intriguing one. If you'd want me to dig into how people are really using AI in their personal lives, say so in the comments and I'll write it.

But here's the part that matters. Even at work, the way people use AI is drifting toward the hardest thing to measure. Anthropic tracks how its assistant actually gets used across the economy, and the pattern that's growing isn't automation. It's augmentation, the AI helping a person think through something they were already doing. Coding, analysis, working an idea into shape. That's judgment work. When someone uses AI to pressure-test an argument, sharpen a strategy, or shape a decision, what exactly are you counting? The hours saved? A better call you can't prove you'd have missed? The output is judgment, and judgment doesn't sit still long enough to measure.

And even the part you can measure isn't as clean as it looks. Workday surveyed 3,200 full-time employees at large companies and found that nearly 40% of the time AI saves gets eaten right back up, fixing, rewriting, and double-checking the output. Only 14% of workers walk away with a clear, positive net gain. So the tidy “we saved five hours a week” number is softer than it sounds, because a big share of those hours went to cleaning up after the tool.

Put it together and you have the real problem. Most of what people use AI for can't be counted.

Where it's actually measurable

The honest answer is that calculating the ROI of AI is real in one specific place: anywhere you can put a measurable, quantifiable number on the output, that's it! When the work produces something you can tally at the end of the day, you can run the math honestly and trust the answer.

Follow that rule far enough and you can see where the smart money is actually going. Not AI bolted onto everything, but narrow problems where the payoff is big and countable at the same time. Think fraud detection in a payments flow, where catching one bad transaction is worth real money and you can count exactly how many you caught. That's a very different bet than turning a chatbot loose on “general productivity” and hoping.

So the first move, whether you're investing in AI or leading a team that uses it, is to stop before you greenlight anything and ask: what are we measuring, what is the outcome we're looking to get from implementing AI. And this one is my favorite, is the team trained on AI and what's our strategy for ongoing investment into AI training of our staff. If you can answer that, the investment and your continued use of AI is justified. If you stumble on the question, you need to reconsider how AI actually helps your organization.

Where it isn't, and what to do instead

And here's the real trap, the one that catches smart people. It's not just that most AI work is hard to measure, we covered that. It's that the work with the biggest payoff is the hardest to measure of all.

Researchers at Harvard Business School split AI's value into two kinds, and that split lines up almost exactly with what you can and can't measure. One kind is operational: faster turnaround, lower cost per task, more output per person. You can put a number on all of it. The other kind is strategic: a sharper business decision, a more resilient operation, a move that gets you ahead of a competitor. That's where a lot of the real payoff lives, and it's almost impossible to measure, because a decision like that only happens once. You never get to run the other choice, so you can't prove yours was better.

In one IBM survey, executives were asked about the AI projects they'd shelved. Every one of them had canceled or postponed at least one, because the cost outran the value. So if the biggest value is the hardest to measure, what do you do?

The discipline isn't to invent a return. It's to sort what you can measure from what you can't, before you spend a dollar. I run every AI use case through the same four questions. Call it the AI ROI Measurability Test.

Yes throughout and you can measure it. Then run the actual math, which is the same ROI formula any investment uses:

ROI = the value of what the AI added, minus what it cost, divided by what it cost, times 100 for the percentage.

The only trick is being honest about both halves. And full cost means the license of the LLM plus the implementation and the governance, not just the license price.

Any no across this test, and you're not measuring a return, you're making a judgment bet. That's allowed, and some of the best moves are bets. Just call it a bet and see what happens.

And if someone hands you their own ROI figure, one question sorts most of it: which studies did you use to land on this figure?

Pull it together. The teams that lose with AI won't be the ones who moved too slowly. They'll be the ones who threw money at it because everyone else was, without ever asking what they were trying to get out of it. Following the trend is not a strategy. Knowing what you're measuring is.

So before your next AI investment, sit with the one question that sorts this out: what are we measuring, what is the outcome we're looking to get from implementing AI?

Has this been your experience with trying to calculate the ROI of AI? Drop it in the comments and let's talk about it!

Sources

  • Peng, Kalliamvakou, Cihon & Demirer, Evidence from GitHub Copilot (2023) — the 55.8%-faster RCT. arxiv.org

  • METR, Early-2025 AI & Experienced OSS Developer Productivity (2025) — the 19%-slower RCT. metr.org

  • Brynjolfsson, Li & Raymond, Generative AI at Work, NBER (2023) — +34% for newest agents. nber.org

  • Pew Research Center, Americans and AI 2026 — self-reported usage mix. pewresearch.org

  • Anthropic, Economic Index (2026) — augmentation over automation. anthropic.com

  • Workday, Companies Are Leaving AI Gains on the Table (2026) — ~40% lost to rework; 14% net gain. workday.com

  • Harvard Business School Online — operational vs. strategic value. hbs.edu

  • IBM — cost-driven cancellation of genAI initiatives. ibm.com

  • Coherent Solutions — source of the Microsoft-commissioned 3.5x figure (vendor-favorable). coherentsolutions.com

Reply

Avatar

or to participate

Keep Reading