Reading Time: 4 mins

The End of Tokenmaxxing: Why More AI Doesn’t Automatically Create More Value

Authored by

One quarter, the dashboard shows rising AI usage, and everyone nods contentedly. The curve climbs, adoption looks solid, transformation is underway. The next quarter, the CFO sits in the same meeting and asks an uncomfortably simple question: What exactly has this additional usage actually improved? Silence follows. The gap between these two slides contains no missing token count. What’s missing is the business denominator against which you could measure consumption at all.

Amazon measures adoption, SAP caps capacity

Two major companies have publicly made two opposite mistakes in quick succession, and both are instructive. Amazon introduced an internal leaderboard that rewarded the use of AI tools. Those who used a lot ranked well. The problem with such a leaderboard is not its intent, but its metric: it measures an activity and treats it like an outcome. Unsurprisingly, employees optimized for exactly what was measured. According to Financial Times reporting on Amazon’s internal AI leaderboard, the list was scrapped once the perverse incentives became visible.

SAP steers from the opposite direction. According to an internal presentation reviewed by Handelsblatt on SAP’s cost-capping approach, the company manages internal AI consumption through a tiered budget model. CFO Dominik Asam is quoted there with the goal of deliberately limiting artificial intelligence costs rather than letting them grow unchecked. This is the mirror image of Amazon’s stance: not celebrating usage, but rationing capacity.

Case Control Metric Intended Effect Blind Spot
Amazon Leaderboard Maximize adoption Roll out AI tools broadly and quickly Usage becomes an end in itself, independent of results
SAP Budget Model Cap consumption Keep costs predictable and tight A worthwhile use case can fail due to budget constraints; a worthless one stays within limits

You could read these two cases as caricatures, but that would be too convenient. Both control metrics are reasonable in their own right. Adoption is necessary, because a tool no one uses creates no value. Capacity is scarce, because compute time and model access cost real money. The shared logical error lies elsewhere: neither high consumption nor low costs prove by themselves that a business ultimately performs better. Adoption and capacity answer different questions, and neither of them is: Does this pay off?

A token bill without reciprocal value is just accounting

Take a concrete scenario. An AI assistant summarizes an incoming customer inquiry before a person handles it. The token price of that summary is visible down to decimal places. It appears in the invoice, you can calculate it per request, it feels precise. That very precision is the trap. The price of the summary is known; its value only reveals itself in comparison with what would happen without it.

This comparison has multiple dimensions. Does the summary actually save the handler time, or does she read the full request anyway for safety? Does response quality improve, or does the model occasionally produce a plausible but incorrect condensation that someone later labors to correct? Does it require human review, and how expensive is that review relative to the saved minutes? Only when these questions are answered does a visible token price become a credible value contribution, or simply an expensive nothing.

The token price is the only number that measures itself effortlessly. That is precisely why it is most dangerous when mistaken for profitability.

From this emerges a simple model that every AI use case should pass. First, a measurable change against a reliable baseline: What was the value before, what is it now, and is the difference attributable to AI. Second, complete costs, not just the token price, but also data preparation, operations, monitoring, and the human labor around it. Third, relevant risks: quality fluctuations, compliance issues, and the cost of a wrong decision based on model output.

For those who want to assess the maturity of their AI deployment systematically, the AI Adoption Maturity Model from Carnegie Mellon Software Engineering Institute provides a governance- and value-oriented framework. The crucial point remains the same: maturity shows not in how much AI a company deploys, but in whether it can name the value of that deployment.

Not every unprofitable attempt is waste

Here the counterargument to the argument itself deserves consideration. If every token must deliver immediate reciprocal value, you suffocate exploration in its infancy. Early experiments produce knowledge first, returns second. A team testing whether a model can sensibly pre-structure offer copy cannot justify that with positive first-month ROI. It buys itself knowledge about an uncertainty, and knowledge has a different price than an ongoing process.

The difference is not whether something pays off immediately, but what you steer by. An experiment needs a learning hypothesis, a small time and cost budget, and a clear endpoint. It asks: What do we want to know, and what can the answer cost? A production application asks differently: What result do we expect sustainably, what full costs does it carry, and when do we observe quality falling below our threshold? Both logics are correct. Yet they are regularly confused, and then an expensive ongoing experiment runs as supposed production without anyone knowing the circuit breaker.

The most expensive AI state is not the failed experiment. It is the experiment that no one ever ended because it eventually became called operations.

The critical move is therefore a conscious transition: an experiment is either converted into production with a value goal and full costs at its end, extended further as a learning initiative, or terminated. This decision matters more than any technical optimization. Model choice, context size, or caching can reduce the cost of a sound use case. For a use case with no identifiable reciprocal value, they only reduce the cost of waste.

Four questions before the next token

The entire argument can be condensed into four questions that should survive before every new and in every existing AI use case. They are no substitute for expert judgment, but they force the silent assumption that usage somehow means value into the light.

  1. What result should change against what reliable baseline?
  2. What full costs arise, including data, operations, and human labor?
  3. What quality, compliance, and decision-error risks are acceptable?
  4. Who decides based on which exit criteria whether to continue, scale, or terminate?

With that, we can return to the dashboard from the beginning, this time with different eyes. The rising usage curve is neither bad news nor good. It is a resource metric that by itself proves neither progress nor waste. The CFO’s question was never unfair. It was just earlier than the answer.

AI maturity ultimately shows not in maximum consumption and not in the lowest budget. It shows in the calm ability to answer these four questions before the curve climbs. Those who can do that no longer measure how many tokens flowed. They measure what is different afterward.

and

Customer Intelligence: Strategies and Use Cases for Leveraging Customer Data

Edited by Prof. Dr. Emanuel Bayer and Manuel Marini · Springer, 2026
The anthology shows how customer data becomes measurable value — with 22 contributions from academia and practice, structured along the five stages of the Customer Intelligence Evolution Framework (CIEF).

  • Framework: Targets & Commitment, Data Integration, Data Quality, Insights, Use Cases
  • Topics: Golden Record, Customer Data Platform, data enrichment, AI-driven personalization, compliance
  • Practice: Field reports from companies including Commerzbank, Deutsche Bank, ADAC, Dun & Bradstreet and IKB

Related Articles

Agentur, iPaaS oder Eigenentwicklung: Drei Wege zur CRM-Integration im Vergleich

Agency, iPaaS or Custom Development: Three Paths to CRM Integration Compared

Reading Time: 7 mins

Agency, iPaaS or build it yourself? Three paths to CRM integration compared, with a focus on costs, risks and mid-market companies.

RevOps ohne Datenintegration: Warum Revenue Operations im Mittelstand an Excel scheitert

RevOps Without Data Integration: Why Revenue Operations Fails in Mid-Market Excel

Reading Time: 5 mins

RevOps promises end-to-end revenue process control. Without data integration, mid-market companies end up with a team that builds Excel spreadsheets.

Bosch hat 400 Tochtergesellschaften: Hierarchische Account-Strukturen im CRM abbilden

Bosch has 400 subsidiaries: Mapping hierarchical account structures in CRM

Reading Time: 8 mins

Corporate clients with hundreds of sites: How machinery manufacturers map hierarchical account structures in CRM, aggregate and bidirectionally synchronize with ERP.