AI demand forecasting involves predicting the future need for artificial intelligence resources, particularly the computational power and “tokens” consumed by AI models. Accurate forecasting is essential for companies investing heavily in AI infrastructure and development, but current methods often misrepresent actual usage and productivity. This can lead to large overestimation of demand, creating a risk of inflated valuations across the AI sector.
Understanding AI Demand Forecasting
AI demand forecasting is the process of estimating how much AI will be used by people and businesses in the future. This includes predicting the consumption of “tokens,” which are the basic units of AI interaction. Every query posed to an AI, every response it generates, and every line of code it writes consumes a certain number of tokens. A simple chat might use a few hundred tokens. However, advanced AI agents, designed to browse the web, write complex code, or complete tasks autonomously, can operate for hours in the background, potentially burning through millions of tokens without direct user oversight.
The challenge in forecasting this demand lies in its variability and the evolving nature of AI applications. It is not simply about counting users but understanding the intensity and efficiency of their AI engagement. Misjudging this can lead to companies building infrastructure for a level of demand that does not truly exist, or failing to account for the actual costs of advanced AI deployments.
The Hidden Costs of AI Consumption
The financial implications of AI usage are proving far greater than many companies initially anticipated. As AI tools become more integrated into daily operations, the costs associated with token consumption are rapidly escalating. Some companies have already seen their AI budgets maxed out surprisingly early in the year. For instance, one major technology company’s chief technology officer noted that their full-year AI coding tool budget was exhausted by April.
Research indicates that companies are overrunning their initial budgets for AI inference, the process of running an AI model, by orders of magnitude. This suggests a widespread underestimation of operational AI costs. Projections even suggest that AI expenses could rival the cost of engineering headcount within companies this year. This rapid acceleration in spending highlights a disconnect between initial budget allocations and the actual rate of AI resource consumption.
The Problem with Current Demand Signals
Several factors contribute to a distorted view of AI demand, making accurate forecasting difficult. One large issue is what has been termed “tokenmaxxing.” This occurs when companies track AI usage as a metric for adoption, sometimes even placing employees on leaderboards based on how much AI they consume. The intention is to encourage AI integration, but it can lead to unintended consequences.
When the goal becomes simply to burn a lot of money, employees may game the system. Instead of focusing on productive output, they might use AI tools excessively or inefficiently to meet usage targets. For example, the CEO of a major AI workload processing company has observed this behavior firsthand among thousands of businesses. Even a prominent CEO in the AI chip sector has expressed alarm if a $500,000 engineer does not consume at least $250,000 worth of tokens, implying a direct link between spending and value. However, this approach risks conflating activity with actual productivity.
Another problem is the inefficient use of powerful AI models. People might use the most advanced and expensive AI models available for simple tasks, such as editing an email, when a less sophisticated and cheaper model would suffice. This over-resourcing for basic functions contributes to unnecessary token consumption and inflated costs.
And, the prevalence of unlimited usage models for consumers has exacerbated the issue. Flat-rate plans, like those offering unlimited messaging for a fixed monthly fee, become unsustainable as AI agents become more common. These agents consume tokens at a much faster rate than simple chatbots. One estimate suggested that a single $200 monthly plan could lead to a burn of $2000 to $5000 in compute costs. Recognizing this unsustainability, some AI labs have begun to cut off popular third-party tools from unlimited subscriptions and are transitioning enterprise customers from flat-rate seats to per-token billing. This shift means a new price tag is attached to every token, forcing consumers and companies to re-evaluate their usage habits and budgets.
The Risk of Inflated Valuations and Overinvestment
The entire AI investment cycle, from chip manufacturing to data center construction, is built on the assumption of continuously growing token demand. Companies selling hundreds of billions of dollars in AI chips, data center operators planning 30 gigawatts of capacity, and major tech firms pouring billions into infrastructure all rely on these demand projections. However, if a large portion of this projected usage comes from employees gaming leaderboards, AI agents running in inefficient loops, or companies blowing through unsustainable budgets, then the underlying demand signal is flawed.
This creates a large risk of overinvestment. Building new data centers, for example, can take one to two years. AI companies are currently making billion-dollar bets on demand that has not yet fully materialized or been accurately verified. This situation is often described as a “cone of uncertainty.” If a company invests too little, it risks losing customers due to insufficient capacity. If it invests too much, the anticipated revenue may not materialize, making the financial math unsustainable. There is little doubt that overinvestment is occurring in some areas. The infrastructure being built today may be sized for a number that is not real, leading to an eventual correction in valuations.
Towards More Accurate Demand Metrics
To mitigate the risks of inflated valuations and overinvestment, a shift towards more accurate and verifiable demand metrics is necessary. Some leading AI labs are already responding to these challenges. For example, one major lab has moved away from flat-rate plans, opting instead for per-token billing. This approach ensures that customers pay for exactly what they use, providing clear data on actual consumption patterns.
This strategy allows companies to build for demand that they can verify, rather than relying on potentially misleading usage signals. The market is beginning to sort this out, with a greater emphasis on understanding the return on investment for AI spending. Setting aggressive goals is still common, but a more prudent and long-term strategy involves underpromising and overdelivering. Companies that can present clean, per-token data, demonstrating what customers are paying for and why, will likely be viewed differently by the market than those that cannot. This focus on efficiency and genuine productivity gains, rather than mere consumption, is vital for ensuring the long-term sustainability and health of the AI sector.