
Loading summary
A
Today on the AI Daily Brief, an operator's cut episode with Nuphar. Everything you need to know about AI Tokens. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, Rackspace Blitzy section and Airtable. To get an ad free version of the show, go to patreon.com americaun or you can subscribe on Apple Podcasts. And to learn more about sponsoring the show, send us a Note@ SponsorsIDailyBrief.AI all right, friends. Well, Nuphar Gaspar is back today and Nuphar and I have been cooking up a lot recently. A whole slew of you have done our most recent program have explored our most recent program, the choose your own adventure style AI summer adventure. Plus we've been cooking up an expanded set of educational resources that we'll be telling you about soon. But one of the realities that both nufar and I have been living in is every company we interact with dealing with the same questions of AI tokens and token economics. We are now firmly in the agentic era of AI, where companies have to think not only about how to get adoption and how to maximize AI's value, but how to do so in a way that doesn't just totally break the bank and where the right intelligence is being used for the right problems. Anyone who's ever built an open clock can tell you this. Getting the right models to do what you want them to without going off into endless cycles of spin takes some real consideration. Today's episode is designed to be the ultimate primer on AI tokens. What we're talking about when we say that term, what the new challenges are and some of the key pitfalls to avoid, as well as strategies to maximize the way that you and your company use AI tokens. All right, Nufar, back with another operator's cut. Talking about the topic du jour, the topic on everyone's minds. We are talking tokens. How are you doing?
B
I'm good. Very psyched to talk about tokens.
A
Yeah, I think it's. I love this period in a discourse where we've gone from sort of pulling hair out, freaking out about the new change, to actually settling into new tactics, new strategy. And I think this is a perfect fit with that. So tell us a little bit about what we're going to be talking about and and let's dive in.
B
Good. So the reason why I wanted to do this episode is because every room that I walk into these days, like literally every room, have some version of the same token conversation. Some practitioners feel like they are being watched when they use an expensive model. The regular users wonder whether one ambitious prompt will eat their weekly allowance. And the leadership teams, they see a bill growing faster than expected. And then they start asking a lot of questions on whether all of these tokens produced anything useful. So I don't know where you guys are sitting, but there is a new anxiety around using too much intelligence. And I actually want to flip the conversation, first of all to make sure that everybody understands what tokens are and what the bill actually means, and then how to spend them wisely rather than sparingly. So that's why I'm here and what I'm planning to do today.
A
One of the places that I've found myself with this conversation is there's been such a visceral reaction now as the cost has gone up. My I have a bigger concern around people retreating back to known ROI biases and not boring, but ultimately low stakes use cases, let's say, as compared to what AI can actually do. That I've found myself in the position of having to defend things like token maxing and token leaderboards just relative to where the tone has shifted. I think obviously we'll get into today the smarter version of that conversation. So I'm excited for it.
B
Exactly. All right. Because there is like a growing conversation around that. I think that there is very like a better language around the feeling of tokens shouldn't be just a financial thing. And Just recently the OpenAI CFO proposed a scorecard and it was called Useful Intelligence for a dollar. And that's built around one question. What does each successful task actually cost? And that's the conversation that I think people should have. And I want to help you with the whole story. What are tokens? What your work costs and where usage creates value and where it quietly leaks value rather than adding. So to kick us off and how we got in here, I want to walk you through four eras of token consumptions and probably will recognize where you are. And we all started by being basically token oblivious. That was the all inclusive era where model companies subsidized the usage and the flat subscriptions hid the emitter and many individual users. And many of them are still there. Just see the ceiling, not a per token price. So that's where we started. And then like you said, we got into the era of token maximizing. That was the leaderboard era, where usage became the badge of AI maturity and we all remember some of the conversations around Meta who tracked employee AI usage on an internal leaderboard. They used to call it a and they used roughly between 60 to 74 trillion tokens in a single month. And according to the data that was published, the top individual user used 280 billion tokens. So to give a sense of how many this is, that is roughly 2.3 million books worth of text. So if you want to try and imagine that, that's about 50 books every minute continuously for a month. So that was the Meta story. And then Uber launched also an adoption leaderboard and burned through the entire 2026 AI coding bud about four months. And there was another company unnamed, but according to TechCrunch they ran up $500 million of cloud bill with no usage limits in place. So then in that era, usage became the metric and dashboard measured activity while claiming to measure value. Obviously this was unsustainable and I know you have some opinions on that, but I'll be curious to say to hear your points. But I also want to say that it actually got us to be as always, the pendulum took us way too far to the era that I call token engine, which is where we are. But if you have anything say in defense of leaderboards, I'm here to listen.
A
Yeah, look, the defense is less Leaderboards are a great concept. They come with I think, a set of very predictable challenges. In fact so predictable that I would say that the hand wringing around the idea of people gaming them has always struck me as a little absurd. Of course people are going to game systems if you put real stakes around them, but that's pretty predictable. And also fairly from first principles you can figure out a lot of ways to deal with that. So I think one, it's overwrought, all the sort of who are freaking out about that. Secondly, my bigger point was that a company that wildly overspends right now via a token leaderboard or anything else, I will bet any amount of money that they will be farther ahead than a company that underspends because they're overly concerned with proving out ROI or whatever it is on a sort of year timescale. Now, the Goldilocks scenario, which I think is what you're going to get into, is being able to experiment, being able to learn, being able to build, and being able to actually understand consumption while also not being afraid of it. But I agree, I, I think that the pendulum swung so aggressively, far too aggressively back from token maxing and excitement to token anxious. And that's the paradigm that we've been living in in the recent few weeks, couple months, whatever it is.
B
Yeah, good. So you, I think you love my model of how to use tokens wisely, but with regards to being token anxious, what I'm seeing in many companies now is that many employees are self censoring themselves, basically trying to avoid costs. And even if we look at the were token maxing. So Meta went from the leaderboard to sending a memo that constrained the AI usage and now the press is calling it token minimizing instead of token maxing. And Uber caps employees at 1500. So even the same companies who are token maxing are now significantly shortening that. And when employees self censor it gets them to feel like every prompt is an ROI conversation and that's not something that we want to have. And I think that this is a very bad era to stay in because I believe that the most expensive token is the one that your best person is afraid to spend. So this is where I want to direct all of us to be in and I call it the token smart era, meaning that you need to spend wisely and not sparingly and understand what creates value and where usage quietly leaks. That's entire kind of theme of this episode and what I wanted to go by in this era, in this discuss, is to talk about the four elements of what a token actually is, why tokens were not born equal, how to audit your own usage and how to govern or do it better within your company. And I'm trying to make it relevant to anybody, whether you are the practitioner that needs to apply some cost engineering playbook and be smart about that, or the executives and admins that need to be proactive and avoid having a difficult conversation with the CFO without a proper response to let's just cut the bill without talking about the business implications, as you just said. So that's the plan for us today. So I want to start with introducing you to the token because I feel that even though it's the most used term in AI, not too many people truly understand what it is because that's in the root of every bill quota and rate limited with AI in tokens. So in a simple word, token is a chunk of a text that the model reads and writes. It's typically bigger than one character and it's usually smaller than a word. And if you have never ever seen a tokenizer or how tokens look in action, OpenAI has a very good page that is open to everyone that you can just take a look at how tokens actually look. So it looks something like that. You can paste the text and then you will see how words are being chunked. So you can see that some words are staying like as one token while others might be separated into multiple tokens. And interestingly numbers often are being chopped in the middle and so on. So we'll put it in the show notes. But a very interesting experiment if you have never seen how your text looks. And by the way, if you paste a non English or a non Latin language, you will see that typically the amount of tokens is much larger than an English language. So that's the OpenAI tokenizer and a few things just to lend it home. In general, the ratio in English is around 3/4 of a word to token, meaning that if you have a page of text it's roughly 1000 tokens. And some languages that are like Hindi, Thai, Greek, other languages like that might get 2 to 5x more tokens for the same content. And because billing is per token, then some questions if you ask them in other languages might cost you much more. And that's sometimes referred to as language tax. With AI with code also it's different and it has its own way. Indentation and brackets and white spaces, they all become tokens. There are some newer ways to tokenize text that are more code friendly in order to do that. But still numbers is a huge problem. So you've seen the 1, 2, 3, 4, 5 being chopped in the middle. And by the way, that's also why whenever everybody's doing like the StrawBerry test for AI and it very badly fails in trying to count how many Rs are in the word strawberry. In many cases that's just a tokenization feature rather than a failure. And the model just has never ever seen the individual letters. It just saw the straw and the berry as separate words. And that's why it's counting it off. So model models have various workarounds, but many of the like AI is so dumb memes are literally just tokenizer issues. So that's tokens in terms of what everyday work costs. I think that's a good kind of mental model to have. So for example, drafting an email is around 500 to 700 tokens. A page of text as noted is about 1000 tokens. You can see longer text. It can be more than that. If you send a model or a tool to do like a AI web search, often it will add a few thousand more tokens for the result, sometimes much more. Images, interestingly are in many cases not that large. In terms of how many tokens there are roughly around slightly more than 1,000 tokens. Interestingly, deep research can very easily be 70,000 or hundreds of thousands of tokens. But just the other day one of our learners in one of our courses had a yes no question and accidentally instead of asking for a web search for the problem, he was asking for the agentic tool to do a deep research. The tool spawned about 100 sub agents to do the research. And then his yes no question cost over 4 million tokens just to answer this question. So it can very easily amount to much more than that. Specifically, a few additional places where you can find very token heavy workloads will be data analysis that can easily get to 1 million or more tokens per task. And heavy coding can also be very aggressive, similarly with many agentic working flows. So just to give you a sense of where it is and to be a little bit more concrete here, the everyday stuff as you've seen, like the emails and so on is almost free. So it's around half a cent. And nobody should ration emails. It's not where the money goes. Search and research can multiply very quietly. So that can be a place to look for efficiency. And the top of the ladder, that's a completely different sport. So if you compare like email to agent decoding task, it can be a factor of a thousand or even more. And another thing that you need to pay attention is that every conversation compounds so the model doesn't remember your previous messages. And as such it sends all of the previous conversations within the same session back to the model. So by, let's say turn number 10, it may be processing so much of the earlier exchange alongside your new message that the total goes much faster than the number of turns suggests. That even happened before the system prompt. And we'll talk about strategies in later on. But this is one of the things that can very easily. Just having very long sessions can very easily amount to to a ton of tokens being consumed.
A
I think this is one of the reasons why this is such an important conversation is another way to put this is that the more advanced and ultimately higher value use cases consume more tokens, which is intuitive that more intelligence is required for bigger challenges. But the direction of use cases is proceeding this way. And so the reason that the token anxiety is going to create problems if not addressed is that it will incentivize people to stay swimming around less sophisticated use cases. So this is the trajectory is clear in terms of less token consumption. The you want as a leader Group your people to be doing more advanced, more useful things with AI. It's just how they do it.
B
Well, so you want them to do deep research where a deep research is required, but you don't want them to accidentally do a deep research on a yes or no question that they can Google in a second. Good. So speaking of the agentic or the more advanced capabilities, those can significantly grow the amount of tokens because agents work autonomously in loops and as such they consume by very widely sighted industry estimates, five to 30 times the tokens of a simple chat. And poorly designed agentic loops or agentic harnesses can be even worse than that because a typical task involves between 10 to 20 model calls carrying instructions and history and tool definition and previous results. And I think According to McKinsey, they estimate that roughly around 60% of an agentic task's cost is tied to the checking and refining and the regeneration of the answers after the first response. So the expensive part is often getting from the answer to the accepted results. So that's an interesting one. And now it gets even more complex because tokens were not born equal. So by the way, the point here is not to not use the agentic tool, just to know that as you said, intelligence cost, but it gets even more complex because tokens were not born equal. And every model lab has its own tokenizer. You've just seen the OpenAI, but different model labs have different tokenizers. So for example, the OpenAI current tokenizer has a vocabulary of about 200,000 tokens, Gemini has around 256,000, Llama by Meta has about half of that, and CLAUDE is unpublished. And the reason why we all should care is that the price per million tokens is denominated in each lab's own tokens and often we don't know them. And the same document can be 10 to 20% more tokens on one provider than another, and even more so for Claude and non English text. So the model behavior widens the gap and one model may answer in a single pass while another reasons longer and writes more and takes more agentic steps and needs retries. And the tool around the model add its own system and context and the loop design. So the two stacks doing the same task can have different token counts and different completion rates. And as a result it's completely different build. So the per token price is kind of the sticker, but the cost per accepted task is the operating metric because otherwise there is no way for you to compare between different providers and different tools. All right, an Important story that also illustrates that what happened when Opus 4.7 came on board the tokenizer basically under the hood changed. And it was this April. And the price sheet was identical to the previous model. The same dollar per million token. But the model was shipped with a new tokenizer that produced by Entropics on documentation they didn't hide it. Roughly 30% more tokens for the same text. So there were quite a few independent analysis of over a million requests that found native tokens they grew and the count grew by about 32 all the way to 45%. And the real world bills grew by 12 to 27% because some of the differ was absorbed by caching. Even Simon Wilson, he measured one of his own prompts at around almost 1 1/2x more tokens. So even though it was documented de facto, we paid more for the same intelligence. And this is like a shrinkflation, right? The same sticker price but a smaller candy bar so nobody prints. Now 30% free awards per dollar, which is the case that happened there. So that's something that is constantly changing. Every lab tuned the tokenizer and often for good reasons. But the operator lessons here is that we have to talk about dollars per task and not dollar per token because the budget is like a moving denominator and it's not the way for you to try and understand how much it's going to cost. Let's talk about what tokens are used for by the AI tools. And you have to understand that every AI request has three token layers and they are priced very differently. We have the input tokens, those will be the prompts and the conversation history and the files and the tools definition and everything that is part of the input. This is what the model reads. And this is the cheapest per token, but can accumulate fast because if the history is being recent or if a lot of context is being read, that can cost quite a lot. Then we have the reasoning tokens. That's the second layer. These are the tokens being used for the model. Internal thinking before answering. For the most part it's going to be invisible to you, but it's billed at the output rates, meaning at the high rate of per token cost. And those can add between 4 to 20 x cost per request. And finally we have the output. That's the answer that you actually see. And this is typically 3 to 5x more expensive than the input price per token. And I think the reasoning layer is the one layer that catches everybody by surprise because you might have a 400 token answer. But under the hood it carried like, I don't know, 4,000 thinking tokens underneath because the model was having an internal monologue and doing a lot of thinking in order to give you the answer. And if you want the analogy, it's like thinking about the part of the restaurant build that is labeled the kitchen time, so you don't get to see it. It's not part of the dish, but you still have to pay a lot of it for that. And the models with the high reasoning effort are often the one with the 20x amount of tokens being consumed versus the lower reasoning efforts. It can be the same question with a significantly different price tag. And sometimes spending a higher reasoning does not get you better results. So some metrics even show that for a simple question it's better to use lower reasoning because the overall cost per task will be significantly lower and the quality will be improved without the model overthink everything. So it's not always that smarter or spending more time thinking gets you better results
A
One of the more interesting shifts in enterprise AI right now is how quickly the conversation is moving towards infrastructure and operations as AI moves into core workflows, regulated data environments, and agentic systems. Enterprises need governed infrastructure and inference that can operate reliably day to day with clear operational accountability built in from the start. As those systems scale, the operating model increasingly becomes part of the AI strategy itself. Rackspace technology is the operator of the full enterprise AI stack, from agents to infrastructure across private cloud, hybrid cloud and edge environments. Rackspace builds and operates governed AI infrastructure, inference and production AI systems for organizations where sovereignty compliance and uptime are non negotiable. Therefore, deployed engineers stay embedded beyond deployment to help operationalize and run AI in live environments. To learn more about where enterprise AI runs and outcomes scale, go to rackspace.com every AI coding tool on the market does the same thing. First it starts writing code. Blitzi does the opposite. Before writing a single line, Blitzi spends days reverse engineering your entire codebase. Thousands of agents ingest millions of lines, mapping every dependency, every undocumented constraint, every architectural decision made over the last decade. The result is a dynamic knowledge graph that understands your software the way a principal engineer would after 30 years in the building. Other tools guess at context with grep searches and markdown files. Blitzi never guesses. It builds true understanding first, then delivers over 80% of entire software epics autonomously validated end to end tested production grade pull requests. That's why Fortune 500 engineering teams trust blitzy with the code bases that matter most. See for yourself@blitzi.com, that's B L I T Z Y.com here's a harsh truth. Your company is probably spending thousands or millions of dollars on AI tools that are being massively underutilized. Half of companies have AI tools, but only 12% use them for business value. Most employees are still using AI to summarize meeting notes if you're the one responsible for AI adoption at your company, you need section Section is a platform that helps you manage AI transformation across your entire organization. It coaches employees on real use cases, tracks who's using AI for business impact, and shows you exactly where AI is and isn't creating value. The result. You go from rolling out tools to driving measurable AI value. Your employees move from meeting summaries to solving actual business problems, and you can prove the ROI. Stop guessing. If your AI investment is working, check out section@sectionai.com that's S-E-C-T-I-O-NAI.com this episode of the AI Daily Brief is brought to you by Hyper Agent, where you run fleets of agents your team can manage together. New users get a thousand dollars in inference. Forget local agents and chat workflows waiting on your laptop to be prompted. Hyperagent deploys always on agents in the cloud, doing real work across the tools your team already uses. Marketing's agent turns competitor, moves into landing pages. Sales agent enriches leads, drafts emails and updates. The CRM ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at HyperAgent built by the team at Airtable. Claim your $1,000 in inference@hyperagent.com AIDAILY Brief.
B
All right, and I think that the gap between input and output pricing keeps widening at the frontier. If you look at the Fable 5, it's about 10, not about it's $10 per million input tokens and $50 per million with the GPT5.6 solid 6x ratio. So we're seeing the gap even widening. And those effort levels that's also something that highly adds the complexity because these frontier models increasingly letting you dial the reasoning eff effort with higher efforts the reasoning tokens are significantly higher and that's probably the dial that you should even be more mindful of even beyond the models because those can easily cost you 10-12x token increase between high or extra high effort to the low or medium. All right, One last thing here, price tag that experimented came from databricks, because a smarter model might not always be more expensive than a less expensive model. What databricks did is they tested coding agents on real engineering tasks from its own code base. And they were using Sonet 5. It was 1.7 times cheaper per token than Opus 4.8. However, Sonnet cost around $2 per task or 2.09 per task versus 1.94 for Opus. So because Sonnet needed more iterations and more reasoning, had to spend way more tokens to get to the same results overall Opus, which is significantly on paper more expensive model, it was cheaper to operate. Which means that we shouldn't just reach to the cheapest model possible, we need to reach to the right model for the task. And that's not easy to get, but something to be mindful. The other thing that matters to the build is the tool itself. So databricks in the same experiment ran the same model at the same thinking effort through different agent harnesses. And they saw that more than a 2x difference in cost per task with the same quality using different harnesses, just because primarily one tool was feeding the model roughly three times less context than the others and thereby the overall cost was lower. So very difficult build to read and very difficult build to navigate. And I'll try to help you as best I can. So bottom line, we're dealing with cost per task and not tokens because otherwise we will not be able to actually compare apples to apples. And the the cost will include the retries, the review, the every like correction that needs an every additional iteration. And then you need to divide by the number of accepted results. That's your cost per accepted task. That's the metric that should aim for and optimize for. And this also brings the conversation much more into return on investment and business value rather than just having a conversation around tokens. That is very hard as hopefully by now you understand to meter good the practical task that you can do, you need to take between five to 10 representative tasks of what you do, run them through tool model or tool options in order to have a good understanding of in your option space what you should do. Hold the input and the quality bar very constant and compare first past success attempts, human correction and elapsed time and the total cost. The winner is the stack that gets your actual work done reliably. I know it sounds like a lot, but if you have a good taxonomy that was optimized for yourself and you know, for the tasks that you do, which Models overall get you better results, potentially with fewer tokens, or if you can do that for your team or your company, and you will need to do that recurrently because things change quickly, then at least you can teach folks that if you're doing that type of research, the recommended model to get you to the overall best quality and the value per task is the following. And so on. That's the current reality that we live in. Okay, now I want to give you a language on how to look at your own tokens and hopefully using that you will be able to distinguish between the tokens that add value to the ones that not so much. Every token that you or your organization spends, in my opinion is one of three kinds. There are tokens that I call tokens that teach. And this is running in both directions, meaning that you teaching yourself. As NetApp Daniel said before, we don't want to stop the experimentation. So the tokens that include the experimentation, the failed workflows, and let me try these three different ways so I will learn. And so these are tokens that are worth spending because you get learning out of them and you can look at them as tuition. And by the way, those also include what you are teaching AI about yourself. So those will be the identity files and the curated context and the knowledge packs, the memory can be counted as those because. Because teaching your AI who you are, what is your context, and learning what works for you in AI is very critical for you to continue moving forward. They look a little bit like a waste on a dashboard, or a lot like a waste on a dashboard because no deliverable ship. But I claim that these are tokens that you need to defend fearlessly because if you won't defend those, you will very quickly go back to just getting AI's help to draft emails and translate between languages rather than moving towards the workflows that matter. And especially if your company and yourself has a lot of catch up to do on where AI is currently at. So I want the tokens that teach to be defended because those are the things that will move the needle beyond the next category, which I call them, the tokens that produce. So obviously those are the most defensible ones because those are the tokens that used to create work the chips. It can be like the final proposal or the research or the code. Obviously that's the thing that is much easier to show the roi. But lastly, we also have the tokens that should be eliminated. And those are tokens that I call tokens that spin. Those can be machines talking to Themselves or automations that nobody is looking at their output or automations that are running too infrequently. Idle agents, bloated context, misused tools and context using Fable to write an email. So using the wrong model, optimized flows and so on. Those will be activity without sufficient output output. So the token smart move, if I need to summarize, is to kill the tokens that spin to tune the production to make sure that it is cost effective and protect the teaching. And that's the order. Like first go and do the audit on your spin tokens and then do the rest. I have a very embarrassing tokens that spin story which I will share in a minute. But I do want to also note with regards to tokens that teach that a failed experiment is as important as a successful experiment. So you should definitely encourage your employees to fail to try because otherwise the tokens they produce will not yield as much value as possible. So it's embarrassing as an AI expert to talk about it, but my open claw was a chief of staff was because it's currently disabled, a chief of staff that I called Chloe and it was using the Entropic API. And because it was using an API it was like auto renewed renewing all the time. And the bills were sent to a secondary inbox. I wasn't really monitoring them and I was seeing that the charges seem quite high. But because I was getting a ton of value and because I was not paying attention to how frequently I'm getting a new bill, I wasn't noticing. And then early June I was traveling so I was not using my OpenClo at all. And still I see that the bill kept coming. So I was saying like, why am I still getting some bills? So I opened the dashboard only to realize that I spent in two weeks $1,500 on an agent that I was not using. So I opened the dashboard like double clicked and I realized that I had have almost 400 million tokens in and almost zero tokens out. So it was a ratio of almost 3,000 to one from input to output. And that's literally the definition of a machine talking to itself and billing me for like an internal monologue that it was running with itself. Looking further, there was a bunch of cron jobs that the openclock created for itself and it was like a compaction job that ran every 30 minutes on empty sessions. And even worse, like the trend was going up. So I first of all closed my open claw and only to optimize it differently. But if it happens to Me in this setup, it can happen to literally everybody and especially when the credit card is owned by your company and not by yourself, often you will not pay attention because you are not sitting on the billing. Yep.
A
And also the fact that you have a bunch of other things that are working well that you might assume it's those things that are amounting for the cost. So one thing that I wanted to mention with spin is that, that I think a lot of the framing of spin, if people were to pick this up, they might assume that it's only mistakes or errors that produce that spin. But that's not always going to be the case. Like sure, this is sort of a in between example where it wasn't exactly an error because it was doing something that it was meant to, but you weren't really paying attention. So it was doing more of it than it needed to. But I think a lot of times spin will also be just ill defining the parameters for a job that you actually do want. An example of this that I have had is I had an open claw going for a while that was perpetually researching new data sources in AI that could help us figure out where the state of certain adoption metrics was right. Every day there's new studies that come out that measure this or measure that and that tell you about data readiness or systems integration or use cases or whatever. And it's too much to monitor for humans, but agents are really good at it. And so this open claw agent was a researcher that its only job was to on a set schedule based on, on its heartbeat, go out and check for new things. And it was never meant to stop. It was always, it was on a specific schedule, but it basically was this continuous research process that was crawling to the ends of the Internet every day and it ended up just not being valuable enough for the cost, but it was doing what it was supposed to to. And so I think part of the auditing spin is also just figuring out what things have accidentally become spin even if they started in the right area. And I think that's why this idea of auditing I think is a good framework because sometimes it's going to be about just updating or changing a process that was valuable as well as catching mistakes.
B
I agree. And even more, I see many automations that people created because they think they will be useful. Like, oh I, I can't read, I have too many slack messages. Let me just create like a slack miner that runs every hour and reads my entire set of channels and that can easily become $1,000 in tokens that literally do a job that moves the needle for nobody, or the morning brief that you created wholeheartedly with the intention to read it every morning, but for some reason you don't find value and you don't read it. So these are the things that you should definitely audit and kill. And my rule of thumb is if you created an automation and for one or two weeks you have never ever used the output, you should definitely kill because it's a definition of a spin. Or if you are using this automation. But there is an a very bad propos between the value of summarizing all of your slack channels to the bill at the end of the month. That's also something that I consider to be a spin and I think that now that everybody gets cowork or GPT work that even more and more within companies because it's so easy to build these automations and without sufficient literacy about how to effectively use the tokens, people create a ton of these automations that look good on paper but don't look so great on the paper of the bill at the end of the month. All right, so let me give you a list for the suspects for silent token spenders. First of all, it's going to be your idle agents and the over frequent jobs. These are going to be the things that run without any meaningful output or way way too frequently. We also in many cases see automations that nobody uses. So it can be like the weekly report or the dashboard that nobody ever goes to read. Additional thing can be what we refer to often as the pre prompt tax. So anything the model runs and reads before the very first prompt prompt. So those include the always on rules or instructions, the skill definition, the tool definition, and so on. And those can very easily if not properly organized amount to many thousands of tokens each run without you typing a single word. So those amount significantly. Many folks also hold the immortal conversation, meaning that they will continue an endless session that keeps carrying old history and old context, often if even creating a poorer quality. We also in many cases see users never feel filtered data retrieval. So instead of just getting 20 rows from a database, they will pull 500 rows or they will process the entire inbox to look for a specific mail that they know what was the subject line and so on. Many other folks will have the context all over the place, so the agent will have to read through a ton of documentation just to understand what are they talking about and what's the truth here as well as rework loops. So anytime that your agent or Your skill or your just day to day user usage gets you to do more iterations just to get the same result. This is like just more tokens being spent on nothing. So these are the immediate suspects. And I want to show you how you can try to potentially identify whether your system is in a spin situation or that the spin to production ratio of tokens is not well formulated. And the thing here is that not everybody can detect in the same way. Some folks have concrete meter, those will be people who are using the API version of the models, they have the API console that they can use. Or if they are cloud code or cursor users, they have usage View. And of course people with admin privileges, they have an admin dashboard. So if you are one of those, you can do the following things. One thing that you should definitely do is the weekend test. Meaning that if you didn't do anything with AI, but you look at your bill and you see that your bill keeps compounding, you know that there are things that are adding to your value without to your bill without any value. That was what happening to me. Also for very extreme input to output ratio. So agentic work legitimately runs with high ratio. But if you get to a point where it's many thousands to one between input and output, in many cases that's empty loops. In my open claw case it was 2600 to 1, which is ridiculous. And if you see that your spend keep rising while the work or the value that you do stays flat, that's also potentially an indication that you are in a scenario of spin and you need to go and further understand what's the case. However, there, there are many folks that don't have direct meter because they are not using one of these tools or they don't have the admin privileges which is probably most of the regular users for them you should probably use proxies. So just go directly to list all of your automation and the scheduled jobs that you own and ask which one of them added business value last week. If you don't know that's a suspect. Then also watch your quota. And if you're burning through your weekly quota extremely fast, especially if you compare it to other people in your setup or on your in similar roles, that might be that you're doing something wrong there. And if you are in an enterprise plan, your admin do have the view at least of how much you're consuming and also the typically the input output that you can just ask them. And there are many places that you can look. There are specific like slash context and slash usage in cloud code. There is also in application visualization now both in cloud code and in cursor that you can just click on the usage meter and try to understand that. So regardless of what and how visible it is for the you, you should definitely put some caps on how much you spend rather than letting the bill just extend all the time. And put some alerts if there is some kind of a significant jump in how much you consume that can be an indication that something is up in your system. So that's for identifying spin and now the habits that we should all adopt to mind our tokens. These are several things that anybody can do immediately that typically improves the token consumption without reducing the business value value. New task is a new session. This one is an interesting one because we will talk in a minute about also model routers. But at least for now, for the most part be intentional about which model you use for what task. Sometimes it's actually going up to like an opus or even Fable class models because they will get the job done in one iteration and overall reduce the spend. In some other cases it's not doing a web search with Fable but rather going to the Haiku or the lower cost of models right? Size your context. Tell your AI what it needs to know. This is a classical Goldilock. Not too much, not too little, but sufficient such that it will not go into endless internal reasoning token loops just to try and understand what you're talking about. Build reusable capabilities. Often when we're just vibing with our model and trying to use it ad hoc rather than sitting down and creating the skills, creating the proper automation, creating the proper agents. We're just wasting a ton of tokens to re ask the the tools to do something again and again. So saying and building proper systems often is one of the best levers that you have to use their tokens wisely and filter everything that you can tell it in which rows of the table the data exist, in which parts of the project board the data comes for, which slack channels, and so on. The more you point the model to the right place, the better the results that you will get. And lastly, in many cases we start doing the work, we realize that the model is completely off. Maybe it's the wrong model, maybe it's missing something. Don't let it spin, just kill the job early and start again while understanding what you do. And this is one of the cases where looking at the model reasoning will go a long way to understanding that it's completely off in the wrong direction. So I would recommend whenever you send the model to start doing something, especially if it's a significant portion of work, open the thinking to understand what the model is and understanding from the task that you gave it. And if it seems to be off, stop and improve the instructions rather than letting it. So that's The Habits Forever 1, two additional levers that you should consider and some of them are very new. So if you are a cloud code user, you can use the doctor Command. This will basically check not only how much like a past installation stake on your machine, but also how are your token divided, whether you have stale skills, stale tool configuration, whether your overall instructions are overly long or overlapping. So it's a very good command that copy created for us that you can go and execute if you are a Claude user. If you are not a cloud cloud user, you can just have your AI tool investigate what the Dr. Command does and basically recreate it for your own tool. Because it's not like a very complex thing to do. It just audits all of your system for you and gives you a structured report with concrete recommendations of things that you can kill because you haven't run them for a while, or things that are duplicated or stale or contradictory that you can potentially reduce significantly. And with regards to routing, a lot of the industry conversations sit right now around the model routing. And you were just talking, I think today or the other day around some interesting M and A around model routing. Picking the right model is still one of the highest return things that you can do. Even if you are able to use like the cursor, automated router or some of the other solutions that are coming our way. Because it's not always going to be. Even if you have like a router in the background, it's not always going to be be as precise as you knowing which model to use. And in many cases you still don't have in your existing tool a good enough or even an existing router.
A
Yeah, I think that we are very early in figuring out the right patterns around routing. Obviously there are a million solutions coming to market. They're all taking slightly different approaches. You have independent experiments from enterprises who are building their own systems that route between, you know, custom models that they've trained as well as, you know, the premier model like it is, there's no one clear approach yet. And even when there do start to be clear use cases and patterns, they may not fit everyone and every use case. I think it would be entirely unsurprising to me or I Expect that routing norms around certain types of software engineering get solved first because it's more deterministic and clear and you can kind of actually have more sort of verified success or not. I think when it comes to knowledge work tasks more broadly, it's going to be immensely more complicated, especially considering how much of our personal model routing that we do right now is about not what the benchmarks would say on a test, but how we like the particular nature of one type of response versus another for a particular context. So I continue to believe that understanding different model capabilities and having model preferences is still a very high leverage activity and it's going to be for quite some time.
B
I agree. And I think the ultimate test was when GPT5 was automatically routing us and all super users or just like more than occasional users, we were all very frustrated by what we got from the auto mode. I think that's the original test that we want control and we will probably even with a great router for many things will continue to be opinionated and rightfully so. Just for people who are also building their own, obviously they have additional levers like you can create more caching and so on. But still for them it's much the same physics. Like the more control you have, the more you are able to be smart about the way you use the models. That's the additional levers. We talked about tokens that teach, and I think that up until now we were very much focused on things that we can do to reduce the bill. But here I want to fight a good fight and say that we want to protect those tokens because those are in many cases the tokens that you spend in order to get much better return. And it's not just about optimizing the bill to go downwards, but rather to improve also the return that we're getting. And often to improve the return we need to improve the tokens that teach. And we're talking about two ways. Whether it's you teaching yourself, meaning that you run the same task using three different models in order to get to this taste of which models you like for each task, or you try the same task in three different ways until you learn which one works best best, or you experiment with a new tool, or you try a new skill or a new automation and it doesn't work and you try something else. So all of these typically gets you overall to much better results from AI. So those should be protected firstly and also the other side of you teaching AI who you are building the systems, adding more context such that you will get much more personalized results or much more organizational aware results. Those are almost always with direct correlation to how much value you get from AI and so does data. I've seen a study of 20k developers that found that the heaviest AI users were roughly twice as productive in terms of the amount of production code that was shift. So in many cases it's actually becoming much more like a smart exploratory user will get you to better results. So we so to summarize what you need to do in two sides. So for the individual user these are the things that you should definitely do. Go and see whether you have tokens that spin. I'm sure that all of us have those idle automations or maybe some of us have even worse scenarios of the amount of tokens being spinned without any business value. Practice those six habits. You can even put them on a post it and just get yourself to work more effectively with the tokens that you have. Do spend the time to invest in reusable capabilities and improved context that the model can be much more selective and discover the relevant context where it matters. I also want you to audit the things on a schedule meaning regularly go back to the system and see what is now stale or maybe something that was working well has become a stale automation. Maybe you need to improve the context the instructions. Maybe you can remove some of the instructions per the new advice coming from Entropic that the modern models need need fewer instructions, not more. And make sure that you protect the learning budget and as needed go and negotiate that with the people responsible for the budget to make sure that you are not now being reduced to the amount of tokens. That leaves you with very little room for exploration for the organizational side. Make the usage visible and then teach the people because when managers and employees see their own they are much smarter about how they use. But make sure that they are not not being encouraged to spend as little as possible, but to spend smartly and also make sure that the budget is by workload and by individuals. If someone is building skills and context and reusable capabilities for the entire team, they need to get significantly higher budget than the person that just uses the tool as a extended Google and all the time we need to make sure that it's by that you tear it up in some organization and some individuals get significantly higher while others potentially less and not just one size fits all for the entire organization and make sure that everybody listens to something like that or that you do an internal training that teaches people on how to be smart about tokens, but not how to spend as little as possible, but also how to be mindful about the ROI and aiming to use tokens for the things that move the needle for the company. So that's the concrete actions for you and the team. And if you want to be even more token smart smart. So beyond the audit of your own usage, we created for you a token gym that you can go and learn and flex your token smart muscles. And if you want to go even further and to learn how to build and work with AI and agents properly, we do have our existing trainings and the next cohorts start on early September. So we'd love to have you there in the executive catch up or the executive agent leadership that will bring you all the way to be very smart about AI or very smart about agents, depending where. That's it.
A
Awesome. Look, I think that this we're always at the beginning when we're talking about things on this show, but this one is I think, particularly inflection pointy, let's say, to use a word that doesn't exist. We are so clearly just at the beginning of figuring out how to organize the relationship between people and the compute and intelligence that they're going to consume Zoom and it is going to be iterative and messy. Which is why I think so many of these ideas that you presented are shared as frameworks, you know, patterns to explore. Right. It's a set of steps that you can take to try to get a handle on these problems. But every organization at the beginning is going to solve them or not in different ways. So thank you for sharing some starting points and you know, we'll continue to evolve this conversation as as the tools around us change too.
Theme:
The episode, “Everything You Need to Know About AI Tokens,” offers a comprehensive primer on AI token economics and smart strategies for practitioners and organizations. Nathaniel Whittemore (“NLW”) and guest Nuphar Gaspar explore how tokens are used and billed in AI systems, dissect token spending anxieties, identify pitfalls and “spinning” (wasted tokens), and present thoughtful frameworks for maximizing value rather than simply minimizing cost.
“We are now firmly in the agentic era of AI, where companies have to think not only about how to get adoption and maximize value, but how to do so in a way that doesn’t totally break the bank.”
—NLW, [00:32]
Nuphar and NLW dig into the realities modern teams face with AI usage, striving for a nuanced understanding of cost, ROI, and organizational learning in the context of AI tokens.
Widespread Token Anxiety:
Every stakeholder is wrestling with tokens—from practitioners afraid of racking up bills to leaders shocked at soaring costs.
“There is a new anxiety around using too much intelligence… I actually want to flip the conversation… to make sure that everybody understands what tokens are, and what the bill actually means.”
—Nuphar, [02:14]
ROI Versus Innovation:
NLW cautions that fear of costs will stifle bold experimentation and reinforce low-value, “safe” use cases:
“My bigger concern [is] people retreating back to known ROI biases and… ultimately low stakes use cases… I have found myself in the position of having to defend things like token maxing and token leaderboards.”
—NLW, [03:01]
Four Eras of Token Consumption:
Nuphar outlines the “eras” of AI token usage:
“I call it the token smart era, meaning that you need to spend wisely and not sparingly.”
—Nuphar, [07:21]
What Is a Token?
Token Consumption Mental Models:
“The everyday stuff… like emails… is almost free. It’s not where the money goes. Search and research can multiply very quietly.”
—Nuphar, [12:56]
Compound Conversations:
“Even though it was documented, de facto, we paid more for the same intelligence. And this is like a shrinkflation, right? The same sticker price but a smaller candy bar.”
—Nuphar, [18:22]
Input Tokens (cheapest):
Prompts, instructions, context—all that the model reads.
Reasoning Tokens (mid):
“Invisible” internal model thinking. Billed more expensively; can be 4x–20x more expensive than input.
Output Tokens (most expensive):
What you actually see and receive. Priced at premium.
“You might have a 400 token answer. But under the hood it carried… 4,000 thinking tokens underneath.”
—Nuphar, [18:52]
Importance of ‘Cost per Accepted Task’:
Model choice, tool design, retry rates, and the complete workflow—not just per-token pricing—determine cost efficiency.
“We shouldn’t just reach to the cheapest model possible, we need to reach to the right model for the task.”
—Nuphar, [23:51]
Three Token Categories:
Memorable Spin Stories:
“Almost 400 million tokens in and almost zero tokens out… that’s literally the definition of a machine talking to itself.”
—Nuphar, [31:16]
Audit and Control Practices:
For Individuals and Teams:
Audit for Silent Spenders:
Six Habits:
Audit Regularly:
Protect the “Learning Budget:”
For Organizations:
Usage Transparency:
Budget by Value:
Training and Literacy:
On cost anxiety and value:
“The most expensive token is the one that your best person is afraid to spend.”
—Nuphar, [07:08]
On agentic workflows:
“According to McKinsey… 60% of an agentic task’s cost is tied to checking and refining and the regeneration of the answers after the first response.”
—Nuphar, [15:53]
On “tokens that teach”:
“A failed experiment is as important as a successful experiment. So you should definitely encourage your employees to fail to try.”
—Nuphar, [28:23]
On audit best practices:
“If you created an automation and for one or two weeks you have never ever used the output, you should definitely kill it because it’s a definition of a spin.”
—Nuphar, [34:00]
On the future of routing:
“I continue to believe that understanding different model capabilities and having model preferences is still a very high leverage activity and it’s going to be for quite some time.”
—NLW, [44:08]
On iterative experimentation era:
“We are so clearly just at the beginning of figuring out how to organize the relationship between people and the compute and intelligence that they’re going to consume… It is going to be iterative and messy.”
—NLW, [49:40]
The conversation is pragmatic, candid, and rooted in real operational anecdotes. Nuphar brings a data-driven, systematic operator’s lens, peppered with direct “lessons learned” stories and practical frameworks. NLW supplements with strategic concerns about culture and innovation, anchoring the discussion in the present-day realities of AI adoption.
This summary covers all critical topics, offers frameworks and concrete behaviors, and highlights memorable moments and speaker wisdom—serving as a rich primer for those new to token economics or seeking to level up how their teams interact with AI in the “agentic era.”