Beyond Tokenmaxxing: the rising token tax on enterprise AI

A person typing on a laptop and using a tablet. Only their upper torso, arms and hands are visible. Text superimposed on the image shows AI
(Image credit: Getty Images)

During the first wave of AI adoption, much of that cost was effectively hidden from customers. AI came wrapped in subsidized pricing, generous allowances and a relentless focus on driving usage. The message was simple: use more AI.

In some organizations, usage itself has become the goal. The rise of concepts like "tokenmaxxing" took this to an extreme, celebrating volume over value and falsely equating outputs with outcomes.

Peter van der Putten

Director of the AI Lab and Lead Scientist at Pegasystems.

But while tokenmaxxing may be a questionable habit that’s been widely exposed, it is not the biggest problem facing enterprise AI.

Latest Videos FromTechRadar

The bigger issue is the hidden token tax that comes with their use of gen AI and agentic AI capabilities in everyday operations.

The token tax is kicking in

Enterprises are already beginning to see the impact. Uber reportedly exhausted its planned 2026 AI budget by April, just four months into the year, after rapid adoption of AI coding tools across its engineering organization.

Amazon reportedly shut down an internal AI-usage leaderboard after concerns that it encouraged “tokenmaxxing,” with executives urging employees not to use AI merely to increase usage metrics.

As AI moves from experimentation to production, organizations are discovering that the economics of scale can look very different from the economics of experimentation.

Tokenmaxxing may encourage organizations to consume more AI, but the token tax is the bill that eventually arrives. And for many enterprises, that bill is proving far larger than expected.

Why is this happening?

On the surface, it seems counterintuitive. Per-token pricing continues to fall, and model providers regularly announce cheaper rates. Yet with enterprise AI bills continuing to rise, the reason lies in how modern agentic systems actually work.

The visible output you receive is only a small part of what is happening behind the scenes. Before generating an answer, an AI agent may interpret the request, decide which tools to use, retrieve data, evaluate results, and loop, calling tools repeatedly until the task is completed or the agent determines it is stuck.

Think of it as an agent having an inner monologue: interpreting the request, planning, calling tools, checking results, and deciding what to do next. Each step may trigger additional model calls and may carry forward more context, so the visible answer can represent only a fraction of the total tokens consumed.

So, while the cost per token is dropping, the number of tokens required to complete a task is often increasing dramatically.

Goldman Sachs Research estimates a 24-fold increase in token consumption by 2030, reaching around 120 quadrillion tokens per month as consumers and enterprises adopt agentic technology.

With this, organizations may believe they are benefiting from lower pricing while their overall costs continue to rise. And the issue is not just cost; it is also predictability, and the lack of a clear link between token usage and outcomes.

This is becoming one of the biggest economic challenges facing enterprise AI, and it is only beginning to be discussed. The economics of agentic AI are making it increasingly expensive to run agents at scale.

What is the token tax?

We position the token tax as the hidden cost organizations incur when AI systems repeatedly consume expensive run-time agentic resources to perform work that could have been designed once and reused many times.

Unlike traditional business software, where the cost of execution is largely fixed, many agentic AI systems effectively rethink the same process every time they run. The more complex the workflow, the higher the tax. Many repeatable enterprise tasks do not require, or even allow for open-ended run-time reasoning.

This creates a disconnect between activity and value. Organizations may be consuming millions of tokens, but that does not necessarily translate into better outcomes. In many cases, it simply means paying repeatedly for the same agentic planning process.

Are we using agents in the right place?

Run-time agents are valuable when solving new, unique and underspecified problems, designing workflows, or exploring options. But using expensive reasoning to repeatedly execute the same business process introduces not just cost but also unpredictability. If you ask an agent to re-reason the same question 100 times, you may get 100 different answers.

To compare this to something in the real world, it's a bit like paying a five-star gourmet chef to invent a recipe every time a new order for a meal comes in. The smarter approach is to use agentic AI once to design the optimal workflow, and then use a lighter-weight AI to select the right workflow that executes consistently.

In other words, use the five-star chef's creativity and expensive brain once a season to come up with a new menu and recipes. Then use skillful but much cheaper chefs to prepare the meal following the recipe that has already been designed.

Shift agentic AI left from run-time to design time

This is where the distinction between design-time and run-time becomes critical. Agents can be very useful at design time, where creativity, exploration, and problem-solving create lasting value.

Run-time execution is different. Here, predictability, consistency, governance and cost efficiency matter most. Run-time agents should only be used selectively, workflows should be used for deterministic, repeatable and governed processes.

By using agentic AI to design processes up front, organizations can dramatically reduce the token tax associated with run-time execution while improving reliability.

Sustainable AI is not about eliminating agents, but about being deliberate about where agents deliver value.

The next phase of enterprise AI

This next chapter will be defined less by impressive demos and more by economic discipline. Organizations will need predictable outcomes, predictable costs and a strategy for reducing the token tax embedded within agentic systems. The winners will not necessarily be the organizations that use the most AI. They will be the ones that generate the most business value from every token consumed.

The companies that address this challenge early by optimizing token consumption and architecting for design-time intelligence and run-time efficiency will not just save money, but will build AI systems that are easier to trust, easier to govern, and easier to scale.

As AI moves from experimentation to enterprise reality, the organizations that minimize their token tax while maximizing business outcomes will have a significant competitive advantage.

We've featured the best IT automation software.

This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today.

The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit

TOPICS

Director of the AI Lab and Lead Scientist at Pegasystems.

You must confirm your public display name before commenting

Please logout and then login again, you will then be prompted to enter your display name.