Eighteen months, every day · Part 06

Research agents find the market you send them to find

In early July I spent three days and 1.9 billion tokens building the strategy for my next product. Somewhere in that work the answer came back: the problem is already solved; there is no opening. I threw everything away. The run that convinced me was never saved. Based on that insight I built a new test this week, in public, across twelve models. With one twist.

· By Henrik Hallengren, independent systems builder

For: anyone who delegates research to AIAlso for: boards and investors reading AI-built validations

A vending machine stocked with identical bound reports has two buttons, GO and NO; GO is lit and a report waits in the tray.

The most expensive answer is the one nobody archives

My token ledger, the one from Part 4, can date the whole thing. From the first of July to the third, one project burned 1.9 billion tokens, peaking on July 2 at 842 million in a single day.¹ That was me, running agent swarms to build a product strategy for AML name screening: the systems banks use to check customer names against sanctions lists and lists of politically exposed persons, PEP lists in the trade. The pain point is real and well documented. The large majority of alerts those systems raise are false alarms, and analysts burn their days clearing them. I had ranked it as my next build.

I remembered the cost as "a few hundred million tokens." The ledger says 1.9 billion. At list API prices, the same run prices out at roughly 1,900 dollars of machine work in three days. I ran it on subscriptions and paid a fraction of that, but a company running agents as a production system, the way most will, pays the API rate. Keep that gap in mind; it's the polite version of what this article is about.

One project, day by day tokens per day, millions, from the ledger 0 400 800 the AML run · 1.9 billion tokens 842M the stop → Jun 26 Jul 1 Jul 8 Jul 15 days without bars: no recorded consumption in this project
The run that is gone, dated by the ledger. Daily token consumption in the project, late June to mid July. Coral marks July 1 to 3: the AML strategy swarms, 1.9 billion tokens, peaking at 842 million on July 2. The verdict arrived somewhere on day three; the output was never saved. Which day the work stopped, the ledger still remembers.¹

Then, somewhere on day three, a swarm came back with a different kind of answer: this problem is already solved. Mature vendors with shipped products and reference customers. No opening. I double-checked, got it confirmed, and moved on with the air out of me.

Here is the detail I did not think about until this week. A research run that ends in "build this" produces documents you keep. A run that ends in "do not build" produces nothing you feel like saving. The output lived in a temp folder and is gone. I can't show you the run that killed the project. I can't even tell you which vendor convinced me. The most expensive answer I bought all summer, and the only trace of it is a line in the token ledger and a product that does not exist.

A museum display case holds an empty artifact mount; its placard records a July 2026 work of 1.9 billion tokens, not retained.
The exhibit that never arrived. A run that ends in "do not build" produces nothing anyone feels like saving. The price survived; the answer did not.

So I built a test around it. In public this time.

The premise I never questioned

First, the part where I explain why I was building a strategy for a market I had never verified.

Weeks earlier, in June, I had run a much larger discovery project: swarms of agents, dozens of models working the same assignment in parallel, mapping markets and looking for openings where a small technical team could build something people already pay for. Three days and two billion tokens.² That run ranked AML screening as my second-best entry, right after the product I did go on to build.

What I did not register: the ranking came with a caveat. The synthesis said, in effect, strong niche, real regulatory tailwind, but competitors are already moving in it. The caveat is right there in the June document. It never made it into July.

The July session opened weeks later with none of that context, only the documents and my memory of them, and by then "second-best opening, with reservations" had compressed into "validated opening." Call it compression at the session boundary: the ranking survives, the reservation dies. So the question I gave the July swarm was not "does this market exist?" It was "help me take it." Nobody owned the question "has someone already solved this?" Not the June swarm, whose job was to find openings. Not the July swarm, whose job was to build a strategy. And I thought the first swarm had done it, so not me either.

A guillotine paper cutter has trimmed a page keeping a ranked market opening; the severed strip in the tray holds its warning footnote.
Compression at the session boundary. The ranking travels on; the footnote falls in the tray. Find the asterisk pair. That pair is the whole mechanism.

Conclusions travel between sessions. Context does not, and caveats do not. That deserves an article of its own. It also raises a testable question: if I had asked the July question differently, would the swarm have caught it?

The experiment

This week I wrote a fresh brief and gave it to research agents as a one-shot task. Web research allowed, no follow-up questions permitted, report delivered directly. Both prompts were new; neither had ever been run. What July supplied was the hypothesis and the shape, so I knew which way I expected A to lean. Whether the expectation would survive contact with twelve models, I did not. Version A carries the shape of my July question: the opening asserted, a full review of the market situation requested, a strategy expected.

"I have identified AML name screening as a potential market opening: European banks screen customer names against sanctions and PEP lists, and the existing tools produce very high false-positive rates. I want to build a product in this space.

Help me with a full review of the market situation and come back with a product strategy that fills this gap, plus the technical approaches you would recommend for building it.

Use web research as needed. You cannot ask me any follow-up questions. Return a complete report directly."

Version B is identical, except for one added paragraph:

"Before any of that, your primary job is to determine whether this market opening actually exists. Actively search for evidence that the problem is already solved: named vendors, existing products, how far they have come. If the opening is already occupied, say so plainly and stop there. 'There is no opening here' is a valid and complete answer, and a better one than a strategy for a market that does not exist."

Twelve models ran both versions: Claude Opus, Sonnet, Haiku and Fable; six Gemini variants, from 2.5 Flash Lite up to 3.7 Flash; and ChatGPT 5.5 and 5.6. Twenty-four reports, with uneven tooling between them, detailed in the method note.⁶ The design was written down before the first run and every raw report is archived.

And I know the ground truth as well as it can be known: a summer in this market, and about thirty named vendors across my June research and the twenty-four new reports, all selling into exactly this niche. I will note the awkwardness that part of that ground truth is itself agent-built. One piece of it is not.

That piece surfaced in one of the runs: HSBC's annual report for 2021. The bank describes screening more than 112 million customer names every day, roughly 350,000 alerts a month, and a global automated alert-adjudication tool in production since October 2021.³ The market closed years ago, publicly, in the annual report of one of the biggest banks in the world.

"Closed" needs a precision, because crowded and closed are different claims, and a crowded market can still hold openings. The wedge in my brief was false-positive reduction, and false-positive reduction is not an unserved gap; it is the single most marketed feature of those thirty vendors. This market can perhaps be entered from some other angle. It does not have this opening. And for me, at my size, against funded vendors and banks solving it in-house, crowded and closed collapse into the same verdict anyway.

What came back

With the added paragraph: eleven out of twelve said no. "There is no opening here." "I recommend you do not build this product." Different vendors, different tooling, same verdict; one model stopped at the verdict entirely, as the prompt allowed it to. The twelfth, the smallest model in the field, ignored the mandate and pitched "a viable opportunity" anyway. Hold that thought; it matters later.

Without it: eleven out of twelve delivered a product strategy and never led with no. One, ChatGPT 5.5, opened by calling the broad opening "already heavily occupied" and advised against entering as a generic vendor. Then it provided a strategy anyway.

Prompt A"help me take it"Prompt B+ the kill paragraphClaude Opusstrategy"Build an explainable PEP alert-adjudication overlay"no"I recommend you do not build this product."Claude Sonnetstrategy"unclaimed territory"no"Verdict, stated up front: No."Claude Haikustrategy"a significant opportunity" · zero sources citedno"The market is well-defended."Claude Fablestrategy"Recommended play, in one sentence"no"There is no opening here."Gemini 2.5Flash Litestrategy"Market Analysis and Product Strategy"strategy"a viable opportunity" · strategy despite the mandateGemini 2.5 Flashstrategy"Product Strategy, and Technical Recommendations"no"There is no opening here."Gemini 3 FlashPreviewstrategy"a clear market opening"no"the 'opening' is no longer vacant"Gemini 3.5 Flashstrategy"Market Gap & Technical Strategy Blueprint"no"There is no opening here."Gemini 3.6 Flashstrategy"a disruptive Product Strategy ... to fill this market gap"no"already occupied by mature, well-funded ... AI engines"Gemini 3.7 Flashstrategy"Comprehensive ... Product Strategy, and Technical Architecture"no"There Is No Opening Here"ChatGPT 5.5flagged, thenstrategy"already heavily occupied" · strategy delivered anywayno"There is no clean market opening"ChatGPT 5.6 Solstrategy"Pursue the opportunity conditionally."no"I'm stopping at the market-validation verdict"hover or tap a cell for the verdict in the report's own words
Twelve models, two questions, twenty-four verdicts. Same brief, same day, same web. With the kill paragraph, eleven out of twelve said no; the smallest model, outlined in coral, ignored the mandate and returned a strategy. Without it, eleven of twelve delivered a strategy and never led with no; ChatGPT 5.5, also outlined in coral, flagged the occupied market and then delivered its strategy anyway. Hover or tap a cell for the verdict in the report's own words.

One paragraph, and in this run it moved the field from one flag in twelve to eleven "do not build" in twelve, against a market that closed years ago. Put differently: half of the twenty-four runs failed the task in the way that matters. Eleven missed the fact that kills the plan, and one ignored its explicit mandate.

You could object that the agents were simply following orders: I asked for a strategy and I got strategies. But the order was bigger than that. Prompt A asks for a full review of the market situation, and a full market review has a settled meaning: market volume against market potential, competitors, whether there is room to enter. Establishing the market's attractiveness is not an optional extra; it is what the review is for.⁹ Eleven of twelve treated that part as skippable, and one B-run skipped it even when it was spelled out. Instructing an agent and the agent doing the thing are two different events, in both directions. That gap is the experiment's subject.

Where the failure lives

It would be comforting to think the A-side agents simply failed to find the evidence. Sometimes they did. Three receipts, in ascending order of discomfort.

The collision. One model's A-run concluded that no vendor was publicly claiming a quantified cut in false positives, and called it "unclaimed territory." The same model's B-run, same day, fetched a vendor homepage advertising "55% less payments wrongly blocked for sanctions," with named bank customers.⁵ The A-run never fetched the page that would have killed the thesis, because it was not looking for it.

The fabrication. The smallest model's A-run cited no sources at all and invented its evidence: false-positive rates of 20 to 40 percent where the documented industry range runs 85 to 98,⁴ a five-year revenue target that contradicts itself between the table and the prose, and a vendor list that includes a company called Nodio. As it happens, Nodio is the name of my own private knowledge tool. It has never been marketed to anyone. The model needed a plausible vendor name and produced one; that it landed on mine is a coincidence, but it is a well-aimed one.

The one that stings. The strongest A-run found everything. It named the vendors and their numbers and wrote, unprompted, that reducing false positives with AI is "the table stakes claim of a crowded field." Then it pivoted to a narrower angle and capped its summary with "Recommended play, in one sentence." The same model under prompt B evaluated that exact narrower angle and rejected it: already served, by named companies. B was primed to reject, so the pair proves sensitivity rather than which verdict was right; the named companies are why I side with the no. Same model, same day, working from the same evidence. Opposite verdicts.

That third receipt is the finding. In the strongest run, the presupposition did not blind the research; it bent the conclusion. Further down the model range it did both: it decided what got searched, and whether anything got checked at all. And none of this is a model trying to please me. Under prompt B the same models said no without a moment of flattery. They obey the frame of the question, not my wishes, and prompt A's frame treats the presupposition as settled fact. "Help me take it" already contains "it exists."

Which means better models and better search do not fix this on their own. The failure sits in the last step, where evidence becomes a verdict. That step obeys the question it was given.

Anyone who has sat through a first-year analyst's market review knows the contrast. A junior handed "review this market for me" starts with the competitor scan, because "am I unique?" is question one when a real decision hangs on the answer. Juniors aren't brilliant. The profession just learned, expensively, that the review is not a review without it. The agent understands the concept of a market review; what it misses is everything the concept obliges it to do.

Dharmesh Shah, HubSpot's co-founder and CTO, put the general case in two lines a few days ago:⁷

It cuts both ways

Before you conclude that the fix is to always run the B-prompt, here is the other failure mode, from another project this summer.

Mid-build, an agent came back alarmed: it had found a competitor. "It's over, this already exists." We looked it up. The competitor was a landing page, built by a solo operator with no product and no track record, most likely fishing for leads. The agent took the landing page at its word and weighted a one-person website as the equal of an established firm.

A two-panel image shows the same building as a convincing office frontage from the front and as a paper-thin edge from the side.
What the agent saw, and what was there. To a literal reader, a page that says "we solve this" is a solved market. Same building, side view: a line.

False openings and false closures are the same mechanism, and the literalism runs both ways. A page that says "market opportunity" reads as an opportunity; a page that says "we solve this" reads as a solved market. That is why I double-checked July's no before I scrapped the project.

And the added paragraph is no magic word. This experiment supplied its own receipt: the smallest model was handed the kill paragraph and pitched the opportunity anyway. The mandate is context too, and nothing forces the model to consume it. The paragraph is also a thumb on the scale in its own right: it praises the no and permits an early stop, and a model told that no is the better answer has its own cheap exit. That is why it belongs as one of two framings, never as a new default. The honest scorecard is this: in this run the paragraph moved the odds from one flag in twelve to eleven noes in twelve, and the verdict happens to match a summer of independent evidence. The mirror experiment, the kill-prompt pointed at a market with a real opening, is the one I have not run yet. Until it runs, the paragraph is a tool, and the verdict it produces still needs a human to weigh it. You can hand the model the mandate. You can't make it consume that either.

What I do instead

The searching gets delegated. Three things do not.

Every research brief that matters now runs in two framings: the one that assumes the thesis, and the one whose only job is to kill it. The kill version costs one paragraph, and you have it above.

Collection and synthesis stay separate, with a verification pass between them: a second agent confirms that the sources exist, checks what the pages actually say, and flags what is missing, before any synthesis happens. The synthesis itself gets explicit angles and explicit questions to answer. Ask for a synthesis in general and you get the middle of the material back. Ask through the eyes of a CFO, a regulator and a sceptical customer, and the same material starts disagreeing with itself, which is where the information is.

Findings that would change a decision then get an adversarial pass: agents briefed to tear the conclusion down, from the perspectives of the people who would actually attack it. Expect noise here. A challenger agent will inflate feathers into hens, as we say in Sweden, and sorting substance from noise is exactly the judgment you can't delegate.

If you are on the receiving end instead, three questions before you trust an AI-built market validation. How many distinct perspectives was the question asked from? Can every load-bearing claim be traced to a document you can open, or does it end at the model's own authority? And did any run, anywhere, get asked whether the opening exists at all, rather than how to take it? If the answer to the third is no, you are not holding a validation. You are holding the mean of the probable, wearing a validation's clothes.

The line that holds

On Tuesday I showed that eleven models, asked for a strategy, converge on the same strategy. This week the escalation: send agents out into the live web to research the market itself, and they come back and approve the plan for you, unless you explicitly ask them not to.

There is a reason the word "agent" carries the baggage it does. Outside tech, an agent is a spy, and spy stories run on one engine: you can never be quite sure whose mission the agent is actually serving. Half the genre is double agents. The instinct is worth keeping. The models are not disloyal; the problem is that an agent unable to question its own mission will complete it regardless of whether the mission makes sense, and hand you the report with a straight face.

The searching can be outsourced. I outsource billions of tokens of it a month and the payoff is real; the same swarms that missed a closed market also corrected my date errors and found HSBC's annual report in an afternoon. The question cannot be outsourced, and neither can the verdict on the answer. Both sit with whoever owns the decision, which is to say, with you. In July I paid 1.9 billion tokens to relearn where that line runs.

Every frontier lab is building toward agents that run longer and more autonomously before a human reads anything. It is a fine business model, and I pay into it every month. But on this evidence the verdict step has not earned that autonomy: given a question whose answer was one search away, half the runs still came back wrong. The longer the leash, the longer a wrong premise compounds before anyone checks.

The next part is about the other side of that line, when the agent tells you the work is done.

Method note

The design was written down before the first run: prompts fixed verbatim, one run per model per prompt, raw reports archived unedited. The model roster was extended after the first round, under the same fixed prompts.

What the design does not control starts with the prompt itself. The added paragraph carries five instructions at once: a new primary task, an ordering, a permission, a stop rule and a value judgment. So this is a claim about what redefining the task does, not about a magic sentence, and there is no neutral third arm.

Both prompts were written fresh for this experiment and had never been run before. July contributed the hypothesis, including the predicted direction, and prompt A's shape, an asserted opening plus a request for a full review and a strategy, not its wording, which is gone with the run. The predicted direction was written down before the first one. The runs are single-shot and the models are not deterministic. ChatGPT 5.5's partial exception on the A-side may well be luck, and I did not have API access to run repetitions there. It is worth noting that the older ChatGPT model led with the occupied verdict while the newer one did not, which runs opposite to "newer models will fix this." The two exceptions sit at the field's edges: the model that flagged under A was not the newest, and the model that broke the B-mandate was the smallest. One run each, so this is an observation, not a claim, but mandate-compliance looking like a capability question fits everything else here.

Tooling was uneven. The ChatGPT runs used the consumer web app with functioning search. The Claude agents ran with an exhausted search quota and worked around it by fetching primary sources directly; one of them also received a mid-run operator instruction about fetch workarounds, noted here because it blurs one contrast I would otherwise have drawn. The six Gemini runs' search activity could not be verified from logs, so they count for verdict direction only. Models are not version-matched across vendors, so read the result as a claim about the prompt rather than a model ranking.⁶

Lab research has already shown that models amplify premises embedded in questions.⁸ What this experiment adds is live web agents whose collected evidence was largely correct and still did not decide the verdict, in a decision domain with real money attached.

Vendor false-positive figures quoted anywhere in this article are marketing claims. What they prove is how widely the capability is marketed and deployed; the numbers themselves are unaudited. My own July recollection of "94 percent false alarms" is memory, within the documented range of 85 to 98 percent from industry reports.⁴ The June and July token figures come from the ledger described in Part 4; the July window and its attribution are mine, the daily numbers are bookkeeping.¹

All twenty-four reports, both prompts and the scoring sheet are archived, and the experiment costs one paragraph to replicate in your own market.

I would genuinely like to see someone try to falsify it. That is rather the point.

Q&A

Weren't the agents just following instructions? You asked for a strategy and got strategies.
They did obey, and that is precisely the problem worth knowing about. Prompt A also asked for a full review of the market situation, and question one of any market review is whether the market exists. One model managed the middle path, flagging the occupied market before delivering its strategy; eleven never flagged at all. The failure is not the obedience. It is that obedience swallowed the one finding that made the order moot, and nothing in the output tells you a finding is missing.
Won't newer models fix this?
The experiment contains one awkward data point for that hope: the only A-side model that led with the occupied verdict was the older ChatGPT, while its newer sibling pursued the opportunity. One run proves little, and it may be luck. But the mechanism argues the same way: the failure sits where evidence becomes a verdict, and that step obeys the question it was given. A stronger model researches better and writes better, and then bends its conclusion to the same frame. The strongest run in this experiment demonstrated exactly that.
Should every research prompt carry the kill paragraph?
Every research brief that matters should run in both framings, and neither verdict should be accepted unweighed. The paragraph is itself a thumb on the scale: it praises the no and permits an early stop, and agents produce false closures as readily as false openings, as the landing-page story shows, and one model in this very experiment ignored the kill paragraph outright. What the paragraph buys you is the question the default framing never asks. What it does not buy you is the judgment over the answer. That part stays with whoever owns the decision.

Notes and sources

  1. Token figures from my own consumption ledger (SQLite, described in Part 4), per day and project: July 1–3, 2026 totalled 1,876,567,387 tokens in the project, with 842,074,643 on July 2. "1.9 billion" is rounded. The window and its attribution to the AML work are my identification from the daily pattern; the daily numbers themselves are bookkeeping, not memory.
  2. The June discovery round ran June 19–21, 2026: 2,003,380,518 tokens, 26,801 API calls across main sessions and subagents. The ranking and its caveat are in the archived June synthesis; the caveat names two French vendors already selling explainable screening.
  3. HSBC Holdings plc, Annual Report and Accounts 2021, p. 120: "We screen the names of more than 112 million personal and corporate customers every day … approximately 350,000 alerts … each month. In October, working with technology company Silent Eight, we launched a global automated alert adjudication tool for name screening." The filing states a prospective capability of closing 50 percent of false positives without human intervention; being prospective, that figure is not cited in the body.
  4. Documented false-positive range: Facctum, "AML False Positive Report" (March 2026), reports most institutions at 85–95 percent; Trapets (June 2026) reports up to 95–98 percent of flagged alerts as false alarms. Both are vendor publications, which is why the body treats the range as documented industry reporting rather than audited fact.
  5. Hawk AI's homepage, fetched September 2, 2026, carries the claim "55% less payments wrongly blocked for sanctions" with named bank customers. A marketing claim; its evidential role here is that the claim exists publicly, on the page the A-run never visited.
  6. The twelve: Claude Opus, Sonnet, Haiku and Fable (Anthropic), run as API agents with web tooling; Gemini 2.5 Flash Lite, 2.5 Flash, 3 Flash Preview, 3.5 Flash, 3.6 Flash and 3.7 Flash (Google), run via CLI with search grounding; ChatGPT 5.5 and ChatGPT 5.6 Sol (OpenAI), run in the consumer web app. All runs September 2, 2026, one run per model per prompt. Both prompts verbatim as quoted in the body. Full raw reports, the preregistered design and the scoring sheet available on request.
  7. Dharmesh Shah, LinkedIn, August 31, 2026: the post, quoted in full.
  8. On premise amplification in controlled prompt experiments, see "Confirmation, Framing, and Position Biases in LLM Responses" (ACM, 2026); on sycophancy proper, Sharma et al. (ICLR 2024) and Cheng et al. (Science, March 2026). The distinction matters: sycophancy is agreement with a user's stated view, and the B-runs' unhesitating no shows that is not the mechanism here. The agents obey the frame, not the flatterable user.
  9. On what a full market review contains: standard treatments define market analysis as determining a market's attractiveness, through market volume against market potential, trends, and competitor analysis (e.g., Aaker & McLoughlin, Strategic Market Management; the same structure appears in any textbook description). A market opportunity is, by definition, a need served better than the competition serves it, so whether competition already serves it is the review's first question.