Tokenomics: The Quirky Quest to Charge for AI
If you’ve ever flirted with a free AI chat, you’ve enjoyed a bargain. Behind the curtain, pillars like Microsoft, Google and Anthropic have poured hundreds of billions into training large language models (LLMs). These models use tiny fragments called tokens to translate prompts into answers. Yet the price of a token is constantly shifting, and the number of tokens a user or business burns is growing faster than the cost per unit can keep up.
For big vendors, the fix is simple: offer paid tiers with extra coding or analytics features and splash a price tag on the premium plan. But most businesses – and governments – still struggle to line up a budget when it comes to the secret fee hidden in each token that fires an AI model.
Simon Gooch of Saviynt notes the dilemma: predicting cost over 12–36 months feels “nonsensical” when the value an LLM delivers keeps shifting. The trick of slicing prompts into tokens means a small tweak can yield a wildly different answer. That unpredictability tugs at the heart of price‑setting. GPB research ILBank forecasts token usage to skyrocket to 120 quadrillion tokens a month by 2030, widening the gap between price and volume.
Microsoft has slowed its engineers’ use of third‑party coding tools after the coding token budget burst in under a year. Uber ran a similar story, exhausting its allowed tokens for AI coding in mere months. The cascading effect is clear: when a single prompt can spawn dozens of tokens across multiple agents, the cost starts to grow like an uncontrolled wildfire.
Companies are learning to cope. Oliver King‑Smith of smartR AI suggests that small firms can tuck tokens into flat‑fee personal accounts, though big platforms warn this will stop soon. Rob Steele of iplicit stresses the importance of precise prompts, likening ambiguous instructions to sending someone shopping without a list: “You don’t want to leave a family member to interpret what goes into the basket.”
Venters, a LSE professor, warns that token costs could balloon when AI is scaled for thousands of users, for tasks like testing, security or guard‑rails. The effect is the paradox of “more AI equals more money, but potentially higher results.” Yet firms must still account for those hidden fees when charging customers. Bill Peterson of Sumo Logic, describing new agentic security tools, highlights the lack of a proven pricing strategy: flat hike, outcome‑based or bundled incident charges still feel experimental. He adds that changing token pricing from providers could turn variable cost models into volatile budgets—something most customers hate.
















