What did fifty billion tokens buy?
This quarter, a few hundred euros a month in subscriptions bought me the delivery rate of a fifty-developer team, a knowledge asset that now answers questions for free, and one expensive no that saved me months of building the wrong thing. On the dashboards companies now run, all of it looks identical to waste: fifty billion tokens, a number that reads like a corporate annual AI budget. So this part is an audit, entry by entry: what did it go to, what was the goal, what came out. The total impresses; only the entries inform.
The scoreboard
Readers of Part 1 met two large international companies where token consumption had quietly become a KPI. For this part I sat down with a friend inside one of them and asked how it actually works.¹
He was disarmingly straight about it. What gets measured is spend. Token usage sits in dashboards at senior leadership level, deliberately: outcomes, he said, are very hard to assess, so they use spend as a proxy. Nobody gets shouted at for burning too little. But the number filters down through team meetings as, his phrase, congratulations for the ones who are using it. He volunteered the quiet part himself: he's sure some people are doing things just to be seen to be doing things.
This is not a stupid company. The dashboard is not malice; it is a measurement giving up.
Then I asked the question underneath it all: what would you measure instead? He called it a great question, circled it, and landed where honest people land.
He didn't know.
Neither did I, and I have more skin in this than he does: on the scoreboard as it stands, I'd be unbeatable.
One person, one quarter
Since early May a tracker on my machines has logged every token my systems consume, and the counter passed fifty billion as this was written.² Priced at API list rates, the mix works out to roughly fifty thousand dollars; what I actually paid is about a thousand euros for the quarter, in subscriptions.² That list-price figure, about 16,700 dollars a month, is more than three quarters of the AI-buying companies in Ramp's invoice data spend on tokens in total, and roughly seven times the median company's entire monthly bill.³ This is one person, working, for one quarter. For placement: Anthropic's average for an active Claude Code day is 13 dollars; my working days price out at thirty to fifty times that at list rates, top percent by every anchor I can find, anchors and caveats in the notes.³
A number that size opens the meeting; it answers nothing. So I went to earn an answer to my friend's question the only way I know: audit my own book, entry by entry, three questions per line. What did it go to. What was the goal. What came out. I expected it to be quick and flattering.
It wasn't, quite.
The entry that flags itself
Start where a cost-hunter would. On 11 July the meter read 2.97 billion tokens, 2.55 billion of them in a single project: the deterministic tax engine, mid-rerun of the local-language rebuild readers of Part 2 know.⁴ My personal record, and on any spend dashboard, the day that gets flagged on sight. In the audit it took about a minute: the most expensive day bought the completeness of the asset that every later answer runs on. Line flagged, line explained, auditor feeling clever. If the whole book read like this, it would be a short article.
Two doors out of one sweep
The next entry changed how I read the book. In June I ran a broad survey across finance, insurance and tax, a few hundred million tokens with one goal: find a place where a small deterministic system could beat slow incumbents.⁵ That single sweep opened two doors. One became the tax engine, the flagship of this quarter; the first commit in its repository is dated 20 June. The other was anti-money-laundering screening at banks, where the overwhelming majority of flagged alerts turn out to be innocent.⁵
In July I went through the second door properly: follow-up rounds, a deep run, then a prototype. A couple of billion tokens in, the conclusion was a no. The problem was better served than I had believed, and the opening wasn't there. I threw the prototype away, keeping one lesson: never accept a research synthesis without reading the reports underneath it, the synthesis is always more confident than the reports; the machine's limit in knowledge work gets a part of its own later.⁶
Here is what stopped me. Same sweep, same method, same goal, two neighbouring lines. One became the most valuable asset I own. The other became a couple of billion tokens for a no.
Nothing in the ledger tells them apart, because nothing could have, in advance.
A well-founded no is still one of the cheapest things you can buy; months of my time on that opening would have cost far more than the tokens did. But I had walked in assuming I could sort my entries into good and bad. The two biggest entries in the book were the same bet until the results came in.
The entries you can barely see
The opposite end of the ledger is nearly invisible. The solar installer's chat agent from Part 1 handles an incoming lead end to end for about half a dollar.⁷ A query against the tax engine costs what the server costs, which rounds to nothing, because the engine is deterministic and no language model decides anything at answer time; keeping its knowledge base current runs to a few hundred dollars a year, inside a product whose licences are priced in five to six figures.⁸
The aggregate says the same. Of my fifty billion tokens, 95.6 percent are cache reads: knowledge already on the server, consulted again instead of shipped in fresh each time.⁹ Built the naive way, the full context injected into every call, identical work costs several times more; two companies can buy the same model and get bills an order of magnitude apart; the difference sits in whoever designed the system around it.
One line I paid in cash, because it ran on another vendor's metered API: an indexing experiment on my tooling, receipt just over a hundred dollars, roughly ten times the model's own advance estimate.¹⁰ A model guessing its own cost without sampling the actual data guesses wrong in both directions; remember that the next time a vendor's model prices your deployment.
Two receipts for the same tool
The flight-simulator engineer from Part 3 handed me his own receipt without ceremony: he pays about two hundred euros a month for his subscription, calls it nothing against what you get, and prices the value in people, like having two or three junior developers working for him.¹¹
I wanted the same receipt with harder numbers, so I had it measured: agents went through nine repositories and counted only delivered code with a git record behind it. About 650,000 lines this quarter, a pace that would traditionally take somewhere around fifty developers, and fifty is the low end; the notes carry the table, the baselines and the biases in both directions, because lines of code stay a weak measure even when they're your own.¹² ¹³ ¹⁴
The gap between his receipt and mine, two or three juniors against some fifty, is explained by neither talent nor the model. His domain barely exists in public text, so every instruction must be spelled out by hand; mine is well-documented terrain, and I spend whole days on research and blueprints before a line is generated. The multiple is set by the domain you're in and the method you bring, worth knowing before you budget on someone else's multiple.
The entries I struck
An honest audit ends at the entries that fail it, and in my book that is a whole category. My subscriptions run on weekly quotas, and when the paying work doesn't cover a week's quota, I build something concrete with the rest.
So the ledger contains a summer-school web app for my kids, where chores convert into points and points convert, after negotiation, into candy, with some summer maths smuggled in as a tax. It contains a Portal-style maths game, later translated into Catalan and Spanish. It contains a friend's band site, rebuilt on a CMS that generates static pages, which is more than the band asked for.
Run the three questions on those lines and they collapse politely. The goal was loose at best; the measurable outcome is a chores-to-candy exchange rate that has stabilised into a currency, which is more than some central banks manage; the market value is zero. As business entries they fail, and I strike them without appeal.¹⁵
Not one of them would have been built at API list price. At a hundred dollars an experiment I'd have closed the terminal and gone outside. That is what the subsidy buys, for me and for every employee on an all-you-can-eat licence: trying costs nearly nothing, so it spills into places no business case reaches. Some of that spill is training. Some of it is candy. A spend dashboard books all of it as engagement.
The goalless dashboard
I went in assuming my job was to defend the ledger. I came out having changed my mind about what it is.
A KPI worth the name is a computed value: several measures set against a goal. EBIT is a KPI. A traffic number on its own is noise; a million visitors converting at half a percent is a worse business than 250,000 converting at twenty. Token burn is a measure in singularity: uncategorised it says almost nothing, and categorised it still says nothing about whether the output was any good. My most expensive day and my expensive no look identical on it. The entries that carry the business barely register on it. The candy app registers beautifully.
That is what my friend's company is running: a goalless dashboard, a scoreboard with one column, and the column is spend. It cannot tell a purchase from a leak, and no upgrade will fix that; the distinction was never in the consumption data. An entire industry launched this year to sell better versions, spend consoles, attribution layers, even a standards body for token accounting, and none of it will audit an entry for you.¹⁶ The distinction lives in the goal each line was spent against, and the goal has to exist before the money goes out, or you are not measuring, you are decorating.
How to read a ledger
The protocol, then, briefly; budget season is coming, and the habits a budget encodes outlive it.
Monday morning does not start with pulling consumption reports; you cannot audit a ledger that was written without goals, and most of this year's ledgers were. Start at the other end. Sit down with the people who own your processes, pick one process to redo with this technology, isolate a trial, write the goal down before the spend, learn, then scale. Ask "on what?" before "how much?".¹⁷ And judge building and running separately: the expensive entries tend to build assets, the cheap ones run on them, and a blind cut takes the asset and keeps the activity. Whether you are 20 people or 2,000 changes nothing in this.
One honest complication: much of what a workforce builds today lives in chat threads and app builders like Lovable or Base44, which makes evaluation genuinely hard even with goals. Fine for prototypes, but systems built by people who don't build systems rarely survive contact with production; the trial you isolate should assume that.
And when the board asks who answers when something goes wrong, the reply that has held up in my work is traceability, systems where you can show afterwards what decided what; you design that in at the start, and a chat log is not it.
The scoreboard, revisited
Back to my friend, who is owed an answer. When he circled his own question, he did get somewhere: hours saved, tied to cost through an hourly rate, and ideally projects created. Then he said the honest thing out loud: those numbers are not going to seem as sexy. He's right, and that is the diagnosis. The goalless dashboard survives because the real numbers are unglamorous.
He also gave me the best defence of the scoreboard I have heard, from inside it: early on, when the job is getting a whole workforce literate, letting everybody build and seeing what survives has a logic, survival of the fittest, the strongest wins. I buy it, as a phase, and his phase has already found something countable: a colleague built a marketing tool in-house for a fraction of a vendor's six-figure quote, and the company owns it outright.¹ No spend dashboard will ever show that find.
So here is my answer, one audit late. Measure nothing new until every line has a goal written next to it; a measure in singularity cannot be rescued by a second dashboard. Then read the three columns: what did it go to, what was the goal, what came out. His hours saved and projects created live in the third column, and they mean something only because the second exists. Spend stays on the page, demoted to the cost side of an entry, waiting for its other half.
What's new is the scale. A spend that reads like a corporate programme can now be one motivated person, and a corporate programme can burn the same money producing nothing but attendance. The total cannot tell you which of the two you are funding. The entries can, if you gave them goals.
Three questions to take with you
- Which entry in your AI ledger bought an asset, which bought activity, and could your dashboard tell the difference?
- What was the goal of your most expensive entry, and was it set before or after the money went out?
- If nothing in your ledger ever bought a no, who is asking the questions an answer could kill?
Q&A
- Is fifty billion tokens a lot?
- As a number, yes: against every public anchor it puts me in the top percent of AI-tool users, though well below the corporate leaderboard heroes whose totals are documented as deliberately inflated. As a cost, no: it runs through subscriptions totalling a few hundred euros a month. That gap between list price and subscription price is itself strategic information: the risk of experimenting is currently subsidised, and the rational response is to experiment while that lasts.
- Shouldn't we just give everyone a licence and see what they find?
- That is the drills-in-a-field strategy: a licence for everyone, with no process analysis first, is sending fifty thousand people into a field with a drill each, hoping somebody strikes oil on land nobody surveyed. Somebody might; that is also the business model of a lottery. Even its best defence, mass experimentation as AI literacy, only works as a phase with an end date. Survey first: which processes leak the most time, and where does knowledge live in only one person's head. Then pilots with the people who own those processes, with goals set before the spend. Part 1 of this series is about what happens otherwise.
- Are lines of code a fair measure of what you got?
- No measure is fair on its own. Lines measure delivery volume, and the notes describe exactly how they were counted and what was excluded. What matters in the end is whether the systems run in daily use and whether someone pays for what they do. Volume is simply the part a sceptic can verify from the outside, so it is the part I published.
Notes and sources
- A senior professional at a large international company, from a recorded interview for this series, anonymised; which of Part 1's two companies is deliberately left open. The paraphrases track his words closely: "it's very hard to actually assess outcomes … so I think they're using spend as a proxy"; "congratulations for the ones who are using it"; "I am sure there are some people who are just doing things to try and be seen as doing things". The question "what would you measure instead" was put to him in the same interview; "a great question" and the circling are his. His alternative measures, hours saved tied to cost through an hourly rate and projects created, and his literacy steelman are from the same interview, as is the in-house tool example; the figures in that example are his own rough estimates, not audited.
- Ledger: self-hosted SQLite tracker, recalculated nightly, covering every Anthropic-model session on my machines since 9 May 2026, main sessions and subagents both. Total at the last nightly recalculation 49.9 billion tokens; the counter passes fifty billion as this publishes. Part 1 quoted the June-to-August slice of the same ledger, 45.7 billion. At API list prices the ledger's mix (95.6 percent cache reads, 3.8 percent cache writes, 0.34 percent input, 0.28 percent output) prices out at roughly 50,000 USD at my actual model mix; list prices as published August 2026. What I paid, per the invoices: one Claude Max 20x subscription at 180 euros a month on a VAT-registered account since late May, a second at about 218 euros including VAT from early July, roughly a thousand euros for the quarter in total.
- The company benchmark: Ramp, "How much do AI tokens cost businesses? 2026 spending benchmarks", 8 June 2026 (ramp.com/blog/ai-token-cost-for-businesses), reporting observed AI vendor payments across thousands of businesses in April 2026: median monthly token spend 2,246 USD, 75th percentile 14,843 USD, 90th percentile 73,030 USD. My roughly 50,000 USD a quarter is about 16,700 USD a month, just above the 75th percentile and roughly 7.4 times the median company's monthly spend; the population is companies that already buy AI, and my side of the comparison is list-price equivalence, not an invoice. In headcount: at Ramp's overall median of 46 USD per employee per month (a monthly figure, not an annual one), 16,700 USD a month corresponds to a company of about 360 employees, though the spread is enormous: by Ramp's AI-intensity tiers the same spend corresponds to roughly 40 people at an AI-intensive firm and nearly 600 at a light adopter. Anchors for the individual placement claim: Anthropic's published guidance puts the average Claude Code user at ~13 USD per active day with 90 percent under 30 USD (docs.anthropic.com, August 2026); Meta's 85,000+ employees averaged roughly 0.7 billion tokens per person per 30 days during the spring 2026 "tokenmaxxing" period (The Information), against my ~14 billion; on the community ccusage leaderboard (viberank.app, ~1,100 self-reported developers, identical counting including cache reads) my total spend would sit around the top 50. The named extreme users sit 3 to 60 times above me, and reporting (Pragmatic Engineer, April 2026) documents that several corporate leaderboard totals were deliberately inflated. The thirty-to-fifty multiple prices my days at API list rates against Anthropic's cost figure; both sides are computed from the same token types at the same price list, so the ratio stands, but my side is list-price equivalence, not an invoice. For an organisation, list price is the honest planning number anyway: team plans top out at a 5x premium seat at 90 euros a month plus VAT, while the 20x subscriptions my ledger runs on are not on the business menu at all, and anything you systematise runs through the API at list price. The subscriptions are capped weekly, so all of this is measured with the brakes on; percentiles above the 90th are triangulation rather than measurement, so "top percent" is the defensible claim, and anything deeper in the tail is my judgement, not a measurement. Raw sources saved per file in the research archive.
- Tracker, 11 July 2026: 2.97 billion tokens total, of which 2.55 billion in the tax-engine project across 16,773 API calls. The engine is Pillar Two (pillartwo.hallengrens.com); the knowledge-base gap and its discovery are told in Part 2.
- The June survey covered finance, insurance and tax at a few hundred million tokens; the same run opened the Pillar Two track, whose repository's first commit is dated 20 June. The July deep dive brought the AML line's total to one to two billion tokens, an order-of-magnitude figure from my logs; the project has no separate track in the tracker. Industry benchmark for rule-based transaction monitoring: PwC (US), "From source to surveillance: the hidden risk in AML monitoring system optimization", September 2010, p. 3: "PricewaterhouseCoopers (PwC) analysis indicates that 90 percent to 95 percent of all alerts generated by AML alert engines are false positives." The figure is sixteen years old and concerns rule-based alert engines, which is exactly the class of system discussed; my recollection of the market I studied was 94.
- The asymmetry between coding work and knowledge work, and what it demands of the user, gets a fuller treatment later in this series; Part 3 laid its foundation (the machine recognises language patterns, and there is no test suite for a market conclusion).
- Internal per-lead cost, approximately 0.5 USD at current volumes; my own metric from production logs. The product's pricing is public at autoply.io.
- Annual knowledge-base upkeep estimated at a few hundred dollars in tokens; my judgement from the running change-monitoring, not an invoiced figure. The estimate covers incremental monitoring at section level; structural retakes like the July run are rebuilds and belong to the building column, not the running one. Licence pricing is public at pillartwo.hallengrens.com, from about 47,500 euros a year for the smallest configuration.
- Tracker aggregate over all Anthropic usage since 9 May: cache reads 95.6 percent, cache writes 3.8, input 0.34, output 0.28.
- The indexing run: about 158 million tokens on a metered API, receipt just over one hundred dollars, against the model's own advance estimate of roughly a tenth of that. The estimate-versus-receipt point stands regardless of the exact advance figure, which I did not log.
- From a recorded interview for this series with the engineer quoted, with his permission, in Part 3; his subscription is a Claude Max plan at roughly 200 euros a month.
- Method: nine repositories, measured on git commit-day counts and line counts. Excluded: package managers and node_modules, build artefacts, vendored third-party code (including a full clone of an open-source terminal project), saved raw HTML research sources, and agent worktree copies, which alone inflated one repository by 1.9 million false lines before filtering. The result:
Total 650,940 lines. *The real estate platform and two of the four smaller projects lack a reliable git timeline; their roughly 130,000 lines are counted in the volume figure but excluded from every rate calculation, and the day count for the four smaller projects covers the two that have a clean history. My workflow commits at every finished delivery, so commit days are a fair record of the days a project was actually built.
Repository Building days Delivered code (lines) Agent-orchestration platform 55 246,948 AI toolkit 16 158,271 Real estate platform n/a* 123,186 Tax engine 23 57,310 Knowledge system 6 54,210 Four smaller projects 4* 11,015 - Classic baselines for delivered production code: Brooks, The Mythical Man-Month (1975), ~10 LOC/day on OS/360; Capers Jones, 16 to 38 LOC/day across projects; McConnell, 20 to 125 LOC/day on small projects (10k LOC) down to 1.5 to 25 on very large ones (10M LOC); one independent developer's own twelve-year average, ~50 LOC/day. Collected in Andy Brice, "How much code can a coder code?", Successful Software, 10 Feb 2017. The division behind "around fifty developers": 520,967 lines on the 104 timeline-safe days is roughly 5,000 lines per building day, set against a deliberately generous baseline of 100 lines per developer-day, the upper end of McConnell's small-project data. The source carries its own warning that lines of code are a weak measure of progress; so does this article.
- The biases, both directions. Why fifty may be low: I work on two to four projects in parallel most days, so a "building day" in one repository is rarely a full working day; the counts include only code that survived, not the iterations (some pages were rewritten dozens of times); and the baseline is senior developers, not juniors. Why it may be high: the classic baselines measure whole project lifecycles including meetings, coordination and maintenance, while my figure measures execution days; generated code spends more lines per unit of function, and Jones's figures derive from function-point analysis; a solo builder pays no coordination tax, which was Brooks's actual point about teams. Also relevant: AI-heavy codebases are reported to churn at several times the normal rate, but churned code never enters this count, because only lines that survived to the git record were counted. For balance: the best-known randomised study (METR, July 2025) found experienced developers 19 percent slower with AI on mature codebases they knew well; my setting, greenfield systems built alone, is close to the opposite pole, and both results can be true. Research and scoping days are excluded on both sides: in a traditional team that work belongs to other roles.
- The quota projects are real and in daily use by their intended audiences of one to three people each. The chores-to-points-to-candy exchange rate was settled in bilateral negotiation and is not disclosed here.
- The 2026 spend-tooling wave, examples: Rippling's AI Spend Console (TechCrunch, 2026), Deloitte's CFO "tokenomics" guidance, and a token-accounting standards foundation launched 4 August 2026; raw sources in the research archive.
- The "on what?" framing and both follow-up questions are from my own answer when this interview question was put to me; the follow-ups in full: were these personal accounts for every employee or specific initiatives, what ROI was measured, and were there goals at all.
- The drills image is mine, from this interview; aimed at licence-first strategies adopted without process analysis. It is not an argument against broad access after the ground survey is done.