Loading summary
A
Foreign. Welcome to this week's AI to rli the Big Story episode of our podcast. I'm Ray Reich, founder and CEO of BenchMarket, and I'm joined by my Big Story co host, Peter Buchanan.
B
Yes, I'm Peter Buchanan. I am the managing partner of New Plan. We work with young tech companies trying to get to the next level in a whole bunch of different ways, and most of them are AI.
A
Okay, well, today's topic, very timely, was the focus of the newsletter that we published this morning, and we're recording this on Tuesday, June 2nd. So for everyone out there, if you want to read more details, just go to AI to ROI podcast on Substack and you'll find it. But it's about all these recent stories that show how AI spending and token costs are accelerating sharply while ROI visibility, or I'll even say measurable ROI, is not keeping pace. In fact, Peter, there was a recent ramp data report that, that said that the average monthly AI token spend across its entire customer base, which I don't know the number, but I think it's
B
more than 1000-050000-50000 customers.
A
Yeah, that token spin has risen 13x since January 2025. And I know almost every CFO ISOC2 token costs so far, and the first quarter and a half, almost 2/4 of 26 is raising is rising anywhere from 2 to 4x. But more importantly, only 27% of enterprise executives say that AI investments have met the return on investment expectations. So I thought that's what we're going to talk about in today's episode.
B
You bet. So the cost side of the equation, talking about it there, it's visible and it's growing fast. The other challenge is the value side is basically invisible to most corporate dashboards. So it could be there, but it doesn't show up in any place that anybody measures. So the episode today covers why the measurement problem exists, what it costs enterprises to ignore it, and what the companies are getting right if they're doing it effectively. So we're using ramp data. We're using our podcast interview that comes out next week with Russ Fraden, the CEO of Leridin. We have sources like exponential view, semi analysis, and we have good case studies. So we're backing it all up with evidence today. Ray?
A
Well, to me, it's almost this AI measurement paradox, Peter. And the core argument is, come on, we're getting more productivity out of our individual workers. And I probably talk to 20 to 30 different companies a month and almost everyone says that their X is being more productive or Getting things done quicker, but it's not generating company level roi. So until we can actually translate these individual productivity gains into demonstrable positive impact on the financial reports, there's always going to be questions. I'll give you an example. Right. We know that engineers are submitting a hell of a lot more pull requests. Sometimes I'm seeing, you know, 2x to 3x sales reps are doing more outbound prospecting or delivering more proposals. Analysts are conducting more research and doing it quicker. So that's a signal at the individual productivity level. But as one senior tech executive who's managing 1000 cloud code using engineers told Exponential View, we are seeing the use of AI coding tools where one plus one plus one plus one equals one and a half, not four. Now, I think there's another kind of law at play here, and it's called Parkinson's Law. Peter. And the core premise of Parkinson's Law is that work expands so as to fill available time left, even due to increased productivity. So in the workplace, if a worker now can do something in four hours versus eight hours, they find something to do with that remaining four hours that may or may not be that positive, or increasing productivity or decreasing costs or increasing profitability. So I think that's a real problem. So I don't think it's a motivation problem. I think most companies, individuals using AI are really motivated, I think, and we're going to talk about it. I think it's a management issue, it's a measurement issue, and it's an organizational design problem.
B
Right. So semianalysis calls this phenomenon AI dark output. So AI generates meaningful economic value and it never shows up on the corporate dashboards. Basically make people happy and make people believe that they're making the right investment. So there's an example in our newsletter. So if the cost of developing a legal document drops from $400 to $5, that's completely fantastic. But the savings register as a token expense. So the cost of tokens goes up even though you've saved $395. That's the obvious expense. So the value.
A
Yeah, that's not. Even though on paper they save that expense because they have less time or less resources. But that's not really saving expenses that show up on the income statement. Right?
B
It's not. Doesn't show up at all. So engineers, the enterprises, count these pull requests you mentioned in before that they can't easily tell whether those pull requests improve the product, sped up delivery, reduce maintenance cost. The activity is measurable, but the outcome absolutely is not. So they just see the expense of all those pull requests.
A
Yeah. You mentioned the podcast that I just recorded with Russ Fraden, the founders here at Lerdin, and they build AI productivity measurement software. So he framed the problem to me as a missing measurement infrastructure problem problem, not a availability to data. The data is there, it's just how do you get it and use it? And I will tell you, there's never been a good measurement infrastructure for the return on investment for software. I mean, every company that we sold enterprise software to had to do some type of ROI analysis. If we pay a million dollars to this software, we're going to save $8 million over the next three years. But after they implemented the software, Peter, they never measured it.
B
Oh, absolutely not. If the software is working and people are using it, why would you do that? That would be their response. Right?
A
Because it wasn't an order of magnitude issue. You know, historically, total IT costs software, hardware, infrastructure people, they're only 3.5 to 6% of total revenue. And software spends itself, maybe it was 2, 3, max, 4% of total employee spending. Of course, labor being the biggest. But now we have engineers whose token bill is approaching 50% to 100% of their compensation. So this cost to compensation ratio changes everything about why measurement matters in the AI era. But Peter, I have a question. So if we have this measurement infrastructure challenge, why are companies allowing their AI bills to explode anyway?
B
Right? So it might not be that they're really allowing it as much as it's happening. And they're just catching up now pretty desperately, actually. So let's talk about this. Why these costs are going up. So first of all, there's a major pricing shift in the software industry. So since the start of 2026, major AI vendors have moved enterprise customers from fixed subscriptions to usage based pricing. So that's GitHub, Anthropic, OpenAI, Google, Salesforce, a whole bunch of companies that sell billions of dollars of software. The second. And so enterprises that paid predictable fixed costs that were pretty easy to track. Well, that's all of a sudden, not only is usage increasing, but costs are increasing with that usage. And CFOs are, you know, they're sort of, they're lagging because, you know, it takes a while for the bills to show up on their desk.
A
You know, the other thing I've seen is a lot of AI agents, even including tools like cloud code, you know, they don't just respond to a human prompt in the moment. They create tasks, they spawn sub agents, they run in Loops. So token consumption is not linear. It can be multiplicative today. Give you an example, the Uber use case. Right. I think Uber deployed CLAUDE code to roughly 5,000 engineers just six months ago, back in December 2025. Of course, they budgeted for it, and initially they were like, wow, intrinsic coding feature usage surge 32%, actually, to 32% of engineers by February 2026 to 84% in March. But you know what happened, Peter?
B
They ran out of. They ran out of capacity. They ran out of token budget.
A
Yeah, well, they. Yeah, they. I don't know if they ran out of capacity, because I think cloud kept giving them capacity, but they sure in the hell ran out of budget, the coo. So they consume their entire annual budget by April. And then this just happened last week, and I called a whisper event. It's where one company reportedly spent $500 million in one month on tokens from one vendor. By the way, Peter, I have to tell you, that's. If that's 500 million, it's probably going to show up as 6 billion of AR on someone's.
B
That's right. That's right. Now we know why Anthropic filed so much earlier for their IPO than people expected.
A
Yeah, but, Peter, there's another, I think, dynamic at play here. And that is how for a few months there, we were incenting companies to use AI.
B
We were. So we wanted people to adopt it, to change their workflows. And so this phenomenon, token maxing, came along. So employees at some companies maximize their AI use, not because they needed to, but because they wanted to show to their bosses that they were AI forward. So Intelligence AI surveyed more than 2,000 companies, and they determined that only 18% of token spend on coding translates into ship products that reach actual users. So Larinden had one customer that was spending a pretty good significant monthly sum on an AI agent running on behalf of an employee who already left the company. And it took them a while to notice. So, as Russ Fradin, the CEO, said, normally if you leave a company, I owe Microsoft $100 for an Excel or Microsoft office. But when it's $20,000 a month and your agent is still running, that's a problem. And then you multiply it by all the other agents.
A
And then there's thought leaders out there who are promoting AI, and it's because your business draws a lot of revenue. Probably one of the most famous AI executives is Jensen Huang from Nvidia. And he's been quoted as saying for every $500,000 in developer salaries. Enterprises should plan on about $250,000 in annual token spend. So even though I don't think that kind of two to one ratio, two parts employee compensation, one part token is going to be reality, I do think because of what we said earlier with it only being 3.5 to 6% of budget, that AI is going to become a compensation level compensation scale operating issue. And very few companies have the organizational design or measurement infrastructure to justify that investment today.
B
Or the money.
A
It's real money. The CTO of, of Meta, Andrew Bosworth, in an internal MEMU memo from last week. Actually it was last month. I think nobody should be using AI tools just for the sake of using them. AI motion is not progress and token usage alone is not a measure of impact. This is the same organization that six year, six months ago created a leader dashboard and the concept of token maxing emanated from.
B
Right? Yeah. Meta does a lot of weird things, but I think they were trying to change a particular culture. They had just bought menace. They were pivoting their strategy. I can understand why they did it, but they didn't do it very long because they figured out how much it would cost. So let's move to. All right, so how about a framework for measuring actual AI productivity? So Laredan has this four step framework that gives enterprises a structured path from cost visibility to measuring business impact. In their framework, most enterprises are stuck on gate one, which is the cost layer. Moving through means redefining a lot of things, right?
A
Yeah, Peter, I think I might frame that a little bit differently. I agree with these four kind of phases where it's cost visibility, utilization proficiency and business impact. I think most companies are now getting cost visibility, often in a rear view mirror because they got the invoice and said holy shit. I also think companies are starting to get more into utilization. Not across the board, but at least they see if people logged into their cloud or chatgpt. But nobody's getting into daily or monthly utilization for everyone who has access to a license, which by the way, some recent research I saw, only about 30% of AI licenses are really getting used more than once to twice a month. But when it comes to proficiency, how effective are employees? That almost is never measured or understood. And of course business impact we know about. So most enterprises are stuck on step one. They only can see the cost and the progression through utilization, proficiency and measuring business impact. It's not optional when you're talking about compensation scale type expenses.
B
Right. I think the place to start is to define what Productivity actually is. It's not token consumption, it's not lines of code, it's not pull requests. It's verifiable, high quality output relative to cost. If you're spending $10,000 on something, hopefully you're generating $20,000 in output. A developer who spends $10,000 a month on tokens, but always ships production quality code, that's a productive developer. But one who spent $3,000 and generates code that needs rollback and rework, that's not a productive developer, even though he spent $7,000 less. Most enterprises have done a pretty solid job of pushing employees to experiment with AI. They've done a poor job of measuring the experiments, learning from them and feeding the results back into the organization. So that, you know, we've talked a lot on this podcast about pilots to production and. And that's still a problem.
A
Well P2P, it brings a whole new thing that Napster would be proud of and not peer to peer pilot to production. Okay, now that was a 20 year old reference. So there was an example that Rush shared with me in software development. And you know, they're pulling information from all the coding assistants, GitHub, cursor, cloud code, even Datadog to build a complexity adjusted AI code quality index. So it doesn't just evaluate whether the code was written with AI assistants, but more. And then a PR was submitted, but it actually measures. Did a passcode review? Did it enter production? Did it stay there? By the way, Peter, I saw some additional benchmarking research last month. That said, even though we've doubled the PR as the pull request, the amount of bugs being identified in that code is up 1.78x so twice as much with 1.7x more bugs. Let me give you GTM examples, right? I really don't care if we're using AI to launch more outbound emails or even to send more proposals. I want to know is my pipeline conversion rates improving? Is the speed that an opportunity moves through the pipeline, pipeline velocity getting quicker? Are my win rates higher? Is my bookings output per rep? You know, those are important traceable categories that really show up on the income statement. Now knowledge works a little bit harder, right? So you know, if you're relying on self reported metrics or even workflow telemetry, there's no clean before or after measurements for a lot of those knowledge tasks. That's why over time, personally I think revenue per FTE at the company level will be a great investor proxy for increased productivity due to AI. I also think things like the percentage of departmental cost per revenue or even departmental cost as a percent of opex. You know, what is sales as a percent of opex or what is marketing as a percent of opex? I think those will be coming down dramatically and I. But they need to start measuring them. And then the last thing I will say is total labor costs. Peter they represent about 15 to 30% of revenue for a company today. You know, manufacturing, product companies, retail companies. It's even higher in human capital intensive businesses like professional services or healthcare where labor is 40 to 50% of revenue. I believe that we're going to have to start measuring what is my total percent of cost and operating expenses for labor and for AI agents because I think we're going to see that mix change and those are going to be two very viable levers that we can pull to get my total operating expense as a percent of revenue down from what it is existing currently.
B
Right. So here's the other thing Ray, that's interesting is companies celebrate and really watch their power users. But I think it's really important. We see a lot of tier one consultancies. Industry analysts press they're focused on power users who are driving up costs. But that's not the right population to optimize for because they are already power users. Laradin analyzed glean usage across their customer base. And heavy users in sale or sales organizations were or measurably more productive than light or non users of glean in sales departments. And here's the lesson, okay? The lesson is if you're investing in a tool, understand the distribution of adoption that what, what, what Laredin calls the fat tail. So not the tip of the spear, but all the people who are in the back who need to have their AI productivity increased use that use training to move the low and mid level adopters up the curve. So the goal isn't to make your employees explorers. The goal is to figure out what works from all the experimentation and then tell people what to do and give them guidance and support.
A
Peter reminds me of my GE management training. Hey, the 20% of top performers, motivate them and sent them but don't spend too much time monitoring them. That 70% in the middle, if you can get them more productive and to perform higher, your company is going to be a rocket ship. And then there's that 10 to 20% at the bottom. And we know what we had to do with them. But one of the things that's interesting and everyone's talking about organizational design and organizational structure needs to change So I think you were doing some research on one potential framework that may provide us some valuable lessons.
B
Yes. So most enterprises, they're adding AI to existing processes without redesigning what's called decision rights, which says where things happen in companies, work in process or workflows, how those decisions are actually made and who should make them. So the very excellent small consultancy with the great substack newsletter Exponential View created an analogy to sort of from history to explain how this worked. They took an electrification analogy that explains the approach to really stepping up in AI. And there's three stages. So the first stage in the Exponential view thing is what they call the light bulb stage. And so AI speeds up the individual worker. Those workers are using chatbots and copilots. Most of those enterprises, most enterprises, they've been here since 2023, they adopted early. They have stars that used AI, these great individual contributors. Then stage two, that's what they call the group drive, is attaching AI agents to existing workflows. So the way factories used to attach wired electric motors to old belt and shaft arrangements to make manufacturing work better and labor a little bit less intensive. So most large enterprises are in essence in this attachment phase now. But of course, the unit drive is stage three. And the unit drive says that rebuilding the decision loops in AI so that AI is the first observer figures out what's actually going on, can make some decisions based on rules, not a tool handed to people who still have to make every decision. Because those decision loops are an absolutely huge constraint. And that, that's called congestion. That doesn't mean getting. Well, it means it's not getting a cold, but it's sort of like a business getting a cold. Right, Ray? They get congestion here in stage two.
A
Yeah. Well, I will tell you and you and I spoke, you know, as we were writing the newsletter for this particular episode. And I think for today's era, I've got a different naming convention. I think stage one is individual employee productivity. That's where we've been for the first three years. Then we're going to get into process efficacy, that's making end to end holistic processes more efficient and effective. And then ultimately I think the vision is to become an embedded AI organization where AI is embedded in everything that we do as a company. Most companies are still in that stage one, employee productivity. And we're getting stuck really transitioning to stage two, the process efficacy beyond experiments and some limited production employments. So we have individual outputs and productivity they're piling up, but they're waiting for Organizational design and decision making that was never optimized for the AI era. So I think what we really need to focus on to make progress in this stage two before we can even think about moving to stage three is redesigning current state processes and organizational structures to align, map and optimize what AI can do.
B
Right. And then, then stage three is the, is sort of the, the little bit of the nirvana stage. So here's what would happen when we have this optimization. So first middle managers move from manually routing information between functions to defining the rules under which agents operate and then they own the outcome that the agents produce. So the human's job is basically settling education thresholds and the government scope, putting those things in place, not approving every single output. The labor market data supports this. Employment in the most AI focused sectors, where companies are pretty well into stage two and working hard to get to stage three, employment in those sectors is falling more than the broader economy. But average wages, the people who are still there are going up. So they have fewer people but higher revenues per employee. So junior workers are displaced first as you would expect, entry level analysts, junior developers, associates, because AI replaces a large part of their function. But the volume of work that's left over just doesn't justify having as many of them. So Russ Fraden said the way to succeed is, is to get 30,000 people to row together. So how do you take this back to hr, the learning and development organization, enable the manager level and the mid manager level. How do you move to that fat part of the curve where everybody's rowing together and you're getting to stage three? That's the key issue.
A
Yeah. You know Peter, the thought on less employees is going to be one of the key measurements of AI. People are starting to shoot holes in that. I was listening to Sam Altman being interviewed by CNBC yesterday. He was in Michigan where they're building this huge Stargate, you know, million square foot data center. He goes, I think we got it wrong. He goes, I think we're actually will seeing new employees, we're seeing employees doing more. So I'm not sure we're going to see significant reduction in headcount because of this. And I'm like, wow, we got a problem Houston. Because if we're not going to see headcount reduction and we're going to spend x percent of revenue on tokens. Right, right. That tells me earnings are going to get hit and that's not going to be something Wall street likes. Hey, we always do this. We always have so many great Conversations. And we're coming up to the end of the podcast. Do you mind if I just blow through kind of three examples of what good and bad look like?
B
Yeah, go for it.
A
So Uber is a not so good. You know, we mentioned that before deployed cloud code to 5,000 engineers. By March, they had 84% of engineers using it. And by April, they actually ran out of budget. And the CEO said, and we haven't seen significant improvement measurable by customer benefit or an acceleration of new product features. But then we have the Lowe's example where they built some AI powered tools to give agents more business context across functions like merchandising, promotions, inventory, and customer service. They deployed three of these, and the SVP of customer service at Lowe's said that by having consistent business context tracking across agents, they can see that they're being more effective. They also are tracking employee token as a core operational metric. And they made sure they had the right data architecture in place. So they saw some, some real benefit. And then I think Petroboss was another example, right, Peter?
B
Yeah, yeah. The Brazilian energy company. Yeah.
A
What was their story?
B
So they had, every year they had a very complex tax compliance process. So they fed 150 pages of complex tax regulations and three months of tax data into a generative AI model. And the result was that three weeks later, the tax department identified 120 million in taxes that they did not have to pay. And it was the first time. It also a corollary was the first time in 15 years during tax season that the tax team could actually go home on the weekend.
A
So $120 million is something a CFO could take to the bank.
B
It totally take it to the bank, but it's even better. Okay, so they had a pretty narrow scope. It wasn't every tax issue that they actually had. They built the model, they took some very complex regulations, they made sure it worked, but now they're deploying it basically across every area in their business where they would possibly pay taxes. And they expect to generate a billion dollars in savings. The scale came from getting the first narrow deployment completely Right. And then expanding beyond that with more competence.
A
Okay, let's wrap up with some tactical usable takeaways. Six things. Number one, measure the outcomes, not just activity, because token spending is going to continue to get up to become more expensive. So you want to be able to compare that to outcomes design for the middle of that distribution curve that 70 to 80% who aren't power users, because your top 10, 5 to 10% are already getting great productivity. Let's focus on the masses and give the CFO real ownership. Don't continue to experiment at the department level. Let them buy AI tools and tokens like they used to buy SaaS and get the CFO maybe working in partnership with the cio actually control the budget?
B
You bet. Also start narrow and close the loop like Petrovas did. Define a discrete domain, get verifiable output, figure out what the financials are actually going to be and then expand from there as you move further along with agentic development and putting AI into your workflows. Redesign your decision rights as you deploy with AI. So try to figure out what rules and context you can give AI so that it can make effective decisions so that you don't get congestion in the decision making process and then treat role changes in your company as a program, not a side effect. So displacing junior analytical roles and evolving middle management, they aren't byproducts of AI development, they're outcomes that the business plans for. So the winners. That's how winners move from AI pioneers to AI settlers is infusing AI into everything they do.
A
And I'll just remind the listening audience who's made it to this point with this what Russ Fradin at Lerdin said Use a framework that you move beyond cost, visibility and then utilization to gaining insights and measurements on proficiency and business impact. And I think it's a pretty good way to end today's episode, Peter.
B
It is. Be sure to subscribe to AI the number two roi.substack.com and leave a review of this podcast wherever you listen to your podcast.
A
Okay, see you next week, Peter, and
B
see you next week.
Episode: AI Math Is Not Adding Up – Where is the ROI?
Date: July 14, 2026
Host: Ray Rike
Guest Co-host: Peter Buchanan
This episode dives deep into a growing paradox in enterprise AI: while AI-related spending—especially on tokens—is skyrocketing, clear, measurable ROI remains elusive for most companies. Using data, industry insights, and recent case studies, Ray Rike and Peter Buchanan explore why traditional methods for calculating software ROI fall short in the AI era, highlight the organizational and measurement challenges, and propose frameworks and actionable steps to address the gap. The conversation is practical, fast-paced, and peppered with both industry anecdotes and hard data, aiming to help listeners cut through the AI hype and focus on real enterprise value.
On AI productivity measurement:
“One plus one plus one plus one equals one and a half, not four.”
— Senior Tech Executive via Exponential View, cited by Ray [03:46]
On runaway token costs:
“By April, [Uber] consumed their entire annual budget... It’s where one company reportedly spent $500 million in one month on tokens from one vendor.” — Ray [10:11]
On what not to do:
“AI motion is not progress and token usage alone is not a measure of impact.” — Andrew Bosworth internal memo at Meta [13:01]
| Timestamp | Topic | |------------|-------------------------------------------------------------------| | 01:29 | Data on explosive AI token costs and low ROI satisfaction | | 03:46 | Real-world productivity not matching aggregate business outcomes | | 05:08 | "AI dark output": value is generated but not measured | | 08:10 | Usage-based pricing and surging token bills | | 09:15-10:11| Uber’s AI token budget crisis | | 10:59 | “Token-maxing” employee behavior and hollow incentives | | 13:01 | Meta CTO’s warning: “AI motion is not progress” | | 14:22 | Four-step framework for AI productivity measurement | | 15:42 | Redefining productivity: focus on outcomes, not activity | | 16:54 | Challenges with pilots to production (“P2P”) | | 20:15 | Power users vs. the middle: “fat tail” lesson | | 22:03 | Need for organizational redesign: Exponential View analogy | | 24:17 | Ray’s three stages: productivity → process efficacy → AI-embedded| | 25:37 | Effects on labor force/wages in AI-heavy sectors | | 27:35 | Sam Altman’s headcount prediction: “employees doing more, not less”| | 28:37 | Case studies: Uber (fail), Lowe’s (moderate success), Petrobras (win)| | 31:06 | Tactical takeaways for measuring and managing AI ROI |
— Paraphrased from Russ Fradin’s framework (Leridin CEO), cited by Ray [14:22, 33:02]
"Use a framework that you move beyond cost visibility and then utilization to gaining insights and measurements on proficiency and business impact." — Ray [33:02]
For more insights and resources:
Subscribe to their newsletter at AI to ROI on Substack.
End of Summary