Everyone knows the why. Nobody knows the what.
Established companies are spending real money on AI and getting very little back, and it is rarely the technology's fault. Up close in large organisations this year, and in my own logs, the same three mistakes keep coming round: they measure the wrong thing, they give the job to the wrong person, and they buy the tool before they understand the work. Here is what that looks like, what it costs, and the order that actually works. The fix is cheaper than the mistake.
A large international publisher, spring 2026
The group had said publicly that it intended to lead on AI. Up close, that spring, it looked like this. Several senior people were involved; none had worked hands-on with a language model. One admitted, to their credit, that not much had happened: licences bought from several vendors, nobody quite agreeing which, a few projects started. Another spent our conversation on how a front page should be redesigned with language models, a fine question for a different job. Along the way it came out that the group had already run vector embeddings across its entire archive, a corpus of hundreds of millions of items at the very least.¹ It had cost enough that they would not be redoing it. Asked what the embeddings were used for, the answer was: nothing yet. Asked why that particular approach, the answer amounted to: the investment is made, and it is not up for discussion.
Nothing in that story is unusual, and that is the point. A serious, well-funded organisation had done the expensive thing, embedding the whole archive, before the cheap thing, deciding what question the archive should answer. It had put people without a working understanding of the technology in charge of choosing it. And it counted tools and projects as progress, because nothing in how work was done had changed yet. You see the same three moves in a fifteen-person company: a paid assistant for everyone, the most technical employee “in charge of AI”, no idea which task it was supposed to change.
Three mistakes, and how they feed each other
The wrong measure: tokens consumed
In two large international companies I know through people inside them, one in IT security, one in industrial manufacturing, token consumption has quietly become a KPI. Nobody is paid a bonus for burning more. But you do get a conversation if you burn less, because low usage reads as low willingness to change. The people there find it bizarre. The one honest use of a usage number is as a searchlight for a few months, to see where people reach for the tool; the day it becomes a target it stops telling you anything.² By spring 2026 the habit had a name, “tokenmaxxing”, and by May the backlash was in print: a CTO whose staff were using the models to check the weather, and Uber, which had put its engineers on usage leaderboards and had spent its AI coding budget for the whole of 2026 by April.³
Volume is the easiest number to move. It's also the least informative one. On a heavy working day I go through a billion tokens or more, and 96 percent of them are cached context the model re-reads on every call, which is exactly why the raw count says so little.⁴ I could double it tomorrow by telling a swarm of research agents to loop until they are satisfied. Nothing of value would come out, but the number would look tremendous. Since June my systems have made 244,771 model calls, and nothing in that count can tell me whether a single one of them was worth making.⁴ At list price my usage since June would have cost around 45,000 dollars, spent on tokens by one person.⁵ On a token KPI I would be the star of the department. That should worry anyone using the KPI.
The number that matters is what a token bought. Earlier this month I spent 179 million tokens indexing five large projects into a knowledge layer: summaries and key facts from every file, all searchable. On the tasks that go through it, it uses two thirds fewer tokens and finishes about 40 percent faster, at within a point or two of the same quality.⁶ It is not magic; on a very large codebase that had only been partly summarised, it did worse. On a token KPI that afternoon of indexing looks like waste and every day after looks like slacking. On cost per outcome it is the best purchase I made all summer.
The wrong person: technical, but not in this technology
When a company that is not itself a technology company decides it needs “someone for AI”, it usually turns to the most technical person on the payroll. In a manufacturer, an insurer, a retailer, that person typically runs technology procurement: which systems to buy, which vendors to trust. They have often not built anything hands-on for years, and almost never with this technology. Knowing how an ERP is integrated tells you nothing about where a probabilistic text model earns its place in a claims process. A fair test, if you are in the meeting: ask the person choosing the tool when they last used it themselves for a full working day.
The people who can actually see the value are the ones who know how the work is done and where the friction sits: the process owners, with the builders beside them. Give them a real understanding of what the model can and cannot do and they will find the applications, because they already know the questions. Then give them the mandate. An initiative that has to be sold upward, again and again, to people who understand it less than the person selling it, moves at the speed of the least convinced person in the chain. Mandate removes the chain.
The wrong end: the tool before the process
Start at the tool and you get the other two mistakes for free: you need a number to justify the licence, so you count usage; you need someone to own the licence, so you pick the technical person. The value lives in the processes, in how input becomes output at every desk. Look there first or everything you bought becomes a pilot that never ends.
Iceland
One example from my own work, with numbers, because otherwise this is just a man with opinions. Earlier this year I built a legal knowledge base spanning many jurisdictions. My agents ran their searches in English, the way anyone would by default, and reported back confident that they had found what there was to find. What gave it away was Iceland. One cited court ruling could not be traced back to a first-hand source, and the agent, searching in English, could not find it again. Iceland has the population of a mid-sized city and a language almost nobody outside it writes in; the primary source would only ever exist in Icelandic. Searched for in Icelandic, it was there, and a good deal more came out with it.
The lesson is not about Iceland. An agent will find information, but it will not instinctively find all of it: nobody had told it that a legal source has to be searched in the language of its jurisdiction, so it searched in English, and every country has enough English-language material to make that look complete. The same holds one layer down, in retrieval by meaning.⁷ Search a market in any language other than its own and you quietly retrieve a fraction of what exists, while the agent tells you it did fine.
So I changed one thing: every jurisdiction was searched in its own language. Same structure, same questions, same team of agents.
| English-only search | Source-language search | Change | |
|---|---|---|---|
| Raw sources in the knowledge tree | 321 files (22 MB) | 1,783 files (146 MB) | 5.6× |
| Lines citing case law | 324 | 556 | +72% |
| Lines citing preparatory works and authority guidance | 55 | 202 | 3.7× |
For this job, more was better: the brief was to miss nothing.⁸ No new model, no new budget line: one person understanding how the tool behaves and asking the right question. The publisher had done the expensive half of the same job without the cheap half, and had nothing usable to show for it.
What it looks like when it goes right
It doesn't have to be a large system. For a restaurant group I worked with, a kitchen that made almost everything from scratch, recalculating product costs took the owner two full working weeks each time: purchases came in continuously, prices moved, and one ingredient could hide behind seventeen article numbers. We built a small LLM-assisted flow that mapped article numbers to ingredients and ran the recalculation, with the owner confirming a handful of items rather than doing the work. Two weeks became about forty-five minutes. He had managed one or two repricings a year before, often wrong at that; now he could reprice as often as ingredient prices moved, see the real margin per dish, and tell the floor which three dishes to push this week. The model did the tedious matching. The human kept the judgment.
Or a solar installer outside Stockholm, whose website chat agent answers from the company's own knowledge base, qualifies leads and books meetings straight into the calendar. July, in the middle of the Swedish industrial holiday, was its best month so far: fourteen leads on organic traffic alone, and one meeting booked on a Saturday lunchtime with nobody in the office.⁹ When existing customers started using the same chat for technical problems, which it was never built for, it took the details, promised a callback and escalated to a human. The company asked for a proper support agent, which now exists. Same people, fewer things dropped.
What to do instead
If a chief executive gave me thirty minutes, whether the company has 20 people or 2,000, I wouldn't start with an AI strategy.
Talk to the process owners first. They already know which processes are blunt or fragile, and they will tell you, because nobody has asked. You do not need to map the whole company; most organisations are a bundle of fairly isolated processes.
Pick one, and pilot it beside the systems you already run. Take marketing. Anyone who has compiled a monthly report knows the drill: log in to ten different places, pull numbers in ten formats, glue them together, and only then start the analysis. Half to nine tenths of the time goes on the gluing; the analysis, where a specialist actually earns their keep, gets what is left. That is a pilot; so is any desk where input and output are clear and the middle is drudgery. Show, quickly, that an LLM-enriched layer next to the ERP or the CRM makes that step faster, cheaper or unnecessary. A visible early win is what brings the other process owners with you.
Then train the people on the tools you build. Skip the prompting tricks and teach them what the model is: a prediction engine with a fixed context window, which fills up and drops what you told it; a strong preference for agreeing with you; and, in my daily use, almost no reflex for saying it doesn't know. Teach them to give it the whole situation and the goal, to end a prompt by asking for its reflection rather than its approval, and to decide themselves. And prepare the managers: when a week's work takes a day, the manager who handed out work on Mondays has to hand it out every morning, and nobody trains managers for that.
Then look at what you already pay for. Software sold to many customers is generalised by definition, and the last twenty percent of your process is usually the part it doesn't cover; I have watched a company buy two procurement systems and still have to build the automation around both. Buy what is generic, build the last stretch that is yours. I have replaced systems like that with ones built for the process, for roughly a year's licence and consulting fees, sometimes half, and nothing after that.¹⁰
None of this is about removing people, and it applies to agencies as much as to their clients. Marketing teams stop gluing data into slides and start doing analysis. Ask a model for a strategy to sell chocolate buns and you get the most generic strategy for selling chocolate buns; you never get out more than you put in. A domain expert with a model becomes sharper, and harder to replace.
Why this is urgent
The price I pay today is not the price of the compute. My usage since June cost me about 630 euros in flat-rate subscriptions; at list price it would have cost around 45,000 dollars, more than sixty times as much.¹¹ Companies do not get that deal. Team plans are metered per seat, the enterprise plan bills usage at API rates on top of the seat, and the moment you build the model into a pipeline or a third-party tool you pay per token at list price.¹² I could not do what I do at that price. The flat rates are the subsidised part, and they are already being capped and metered. Behind them sits data-centre spending that has begun to outrun the operating cash of the largest companies on earth, on chips whose useful life is itself in dispute, while the open models trail the frontier by months rather than years.¹³ Once no lab can win on the model itself, no lab has a reason to subsidise your usage to win you over. What is left is paying for what you use.
And when you pay for what you use, what a token bought becomes the only number that matters. The organisations that trained themselves to burn as many tokens as possible will find that habit very expensive. The ones that learned to get the most value from the fewest tokens will not. That is decided now, in what you measure, who you give the mandate to, and which end you start from.
Three questions to ask your own organisation this week: Which process would we pilot first, and who owns it? What did last month's tokens actually buy us? Who here understands both the work and the model, and do they have the mandate to change anything?
Two questions I get asked
- We rolled out a chat assistant to everyone. Isn't that step one?
- It is a purchase, not a step. The research on what happens next is not kind: in a controlled study last year, experienced developers given AI tools took 19 percent longer to finish their tasks while believing they had been faster; in PwC's survey of chief executives this January, 56 percent said AI had brought no significant financial benefit yet, and one in eight had seen both cost and revenue gains; and the much-quoted MIT NANDA report carries a finding that is inconvenient for someone who builds for a living, and true: bought-and-adapted tools succeeded about twice as often as internal builds, mostly because the ones that succeeded were process-specific and judged on outcomes, which is the order this piece argues for. You never get out more than you put in, and if the people holding the tool don't know what to use it on, nothing says their output will improve. The hour a model saves does not turn into money by itself either: in Danish payroll data, only a few percent of the time saved ever reached anyone's pay.
- How do we know if our token spend is reasonable?
- You can't, unless you measure what came out. Not tokens; value. How much faster is the work done, how much more of it gets done, what did it earn or save. If the value out has not risen with the spend, the spend was not reasonable, whatever the dashboard says. Spend 179 million once to cut two thirds of the tokens on every task after and you have made a good investment. Spend more and get no better output, and you have made a bad one. And give it time: in company spending data the effect shows up six to twelve months in, so expect the first months to cost more than they return. The pilot's job is to make that stretch short.
Notes and sources
- “Hundreds of millions” is a floor, not an estimate of the total. Details withheld to keep the company unidentifiable; the conversations were held in confidence.
- Hacker News thread on “Tokenmaxxing is dead, long live tokenmaxxing”, 28 to 29 June 2026 (news.ycombinator.com/item?id=48708795): one commenter describes an employer that tracked usage for a few months explicitly to find where to invest (“throw AI at random parts of your job so we can generate feedback from employees on where to invest in additional automation”) and reports “a ton of high-value little AI workflows” as the result; another argues that “cost-maxing has hurt our company”. Anecdotal, but it is the strongest counter-argument to the KPI critique that I have seen, and it holds only while the number is exploratory.
- Axios, “AI sticker shock hits corporate America”, 28 May 2026 (the CTO and the weather; the term “tokenmaxxing” had been in circulation since March, in the New York Times and TechCrunch among others). Business Insider, 25 May 2026, reporting Uber COO Andrew Macdonald's remark on the “Rapid Response” podcast that AI costs were getting “harder to justify”; The Information, 14 April 2026, reporting Uber CTO Praveen Neppalli Naga saying the AI coding budget for 2026 was spent by April; Forbes, 17 May 2026, on Uber's internal usage leaderboards.
- My own Claude Code logs, in practice 1 June to 18 August 2026 (the log begins 9 May but May is almost empty): 45.7 billion tokens across 244,771 model calls, of which 95.6 percent cache reads, 3.8 percent cache writes, 0.4 percent uncached input and 0.3 percent output. Active on 67 days, 16 of them weekend days, seventeen of them over a billion tokens; three weeks of the period were holiday. Heaviest day 11 July, 2.97 billion tokens.
- Priced at Anthropic's published API list prices, August 2026 (Fable 5 at 10/50, Opus at 5/25, Sonnet at 3/15, Haiku at 1/5 dollars per million tokens; cache writes at 1.25× and cache reads at 0.1× the input price): about 44,900 dollars, of which cache reads 27,500, cache writes 12,400, output 4,000, input 950. With a one-hour cache TTL the figure would be nearer 52,000. Annualised at the same rate, roughly 215,000 dollars, and my usage has been rising. My heaviest single day prices out at around 2,500 dollars.
- Benchmark run 10 April 2026 on a 414-file codebase: ten questions, three models (Haiku 4.5, Sonnet 4.6, Opus 4.6), identical instructions with and without the knowledge layer, answers scored by two separate frontier models (Opus 4.6, with a second opinion from Gemini 2.5 Pro). Tokens per question fell 66 to 76 percent, tool calls 71 to 74 percent, time 37 to 43 percent; quality fell 0.5 to 2.4 points on a hundred-point scale, giving three to four times the quality per token. On a 599-file document project the layer beat the standard tools on all three counts. On a 24,000-file codebase where only a quarter of the files had been summarised, it did not; the approach works best on projects of a few hundred to a couple of thousand files. The layer holds summaries at three levels of detail plus the key facts from every file; keeping it current costs cents to a dollar a day. Indexing the five projects mentioned cost 179 million tokens, mostly on Gemini 2.5 Flash, roughly a hundred dollars.
- Retrieval by meaning (vector search) is language-sensitive too. Same metadata layer built once in English and once in Swedish, queried in Swedish; the Swedish layer matched better on every test question. Legal concepts add a further problem: they rarely translate one to one, and countries package the same law under different names.
- The base was meant as the ground truth for a deterministic legal system, so coverage was the brief. First-pass runs are never the last; a knowledge base like this is challenged through several further rounds and from several angles before it settles.
- Client since 18 May 2026. 22 leads between 1 July and 16 August, 9 of them outside office hours, roughly four in ten; monthly leads 2 (May), 8 (June), 14 (July), 8 (August to the 16th); four meetings booked by the agent in the period, one on Saturday 18 July at 12:30; roughly 45 to 50 percent of chat sessions occur when the office is closed. The company reports no paid traffic to the site.
- One example: a business-intelligence platform built for about 15,000 euros, replacing a licence of roughly 150,000 kronor a year plus a consultant who billed around 250,000 kronor the final year. Approximate figures, from invoices I saw at the time; I no longer have access to check them.
- Two Claude Max subscriptions at 180 euros a month, one of them for a single month of the period.
- Anthropic's published plans, August 2026: a team premium seat gives roughly a quarter of the usage of the top individual plan for half its price, that is twice the price per unit, metered per seat with weekly caps; the enterprise plan is priced per seat plus consumption at API rates; API and third-party integrations are billed per token at list price.
- Alphabet Q2 2026 earnings release, 22 July 2026: net cash provided by operating activities 39.1 billion dollars, purchases of property and equipment 44.9 billion, free cash flow for the quarter minus 5.9 billion; trailing twelve months still positive. On chip life: CNBC, “The question everyone in AI is asking: How long before a GPU depreciates?”, 14 November 2025: Google, Oracle and Microsoft depreciate servers over up to six years, Amazon shortened part of its fleet from six to five years in 2025, and Michael Burry argues the real useful life is two to three. On open models: Epoch AI, “Open models lag state-of-the-art closed models by 4 months”, 29 May 2026 (about six months on a stricter criterion).
- METR, “Measuring the impact of early-2025 AI on experienced open-source developer productivity”, 10 July 2025: developers took 19 percent longer with AI tools while expecting a 24 percent speed-up (16 developers, 246 tasks on codebases they knew well; the authors caution against generalising). PwC, 29th Annual Global CEO Survey, press release January 2026: 12 percent of CEOs say AI has delivered both cost and revenue benefits, 33 percent gains in either, 56 percent “no significant financial benefit to date”. Humlum and Vestergaard, “Large Language Models, Small Labor Market Effects”, NBER Working Paper 33777 (2025), linking adoption surveys of about 25,000 Danish workers to payroll records: time savings of about 2.8 percent of hours, no significant effect on earnings or recorded hours, and only 3 to 7 percent of the productivity gain passing through to pay; BCG's June 2026 “AI at Work” survey adds that 66 percent of regular users get limited or no guidance on what to do with the time saved. MIT NANDA, “The GenAI Divide: State of AI in Business 2025” (a preliminary v0.1 report based on 52 interviews, 153 survey responses and 300 public initiatives, which describes itself as directionally accurate): 95 percent of organisations getting zero return on 30 to 40 billion dollars of enterprise generative-AI investment, with a steep drop-off between pilot and implementation. Two of its findings cut both ways for a builder like me: purchased-and-adapted tools succeeded about twice as often as internal builds, and the pilots that did succeed were process-specific, integrated into workflows and judged on business outcomes rather than usage.
- Ramp Economics Lab, “Heavy AI adopters hire more”, 30 June 2026: among firms on Ramp's card and bill-pay data, the top third by AI spending intensity grew headcount about 10 percent more over two years than light adopters; “firms don't increase their headcount until 6-12 months following adoption”, and the gains are “subject to a learning curve and a minimum threshold of adoption”. Spending data, not outcome data; I use it only for the time horizon.