When in Rome, search in Italian.
We have a saying in Swedish: speak with farmers in the farmers' language. The good thing with an LLM is that you can speak with it however you like and it will still understand you. But that doesn't mean it knows to speak the right language itself. A language model asks its questions in English by default, and most of the answers a European needs were never written in English. I found this out through a legal knowledge base that looked complete and had quietly missed most of two categories of sources. Fixing it took one instruction, a day of processing and just over two billion tokens, but it did multiply the source material five times over. Here is the story of how the gap hides, what it cost to close, and how to keep it closed. The instruction is free; not knowing you need it is not.
Why Iceland gave it away
This spring I built a legal knowledge base covering the global minimum tax across more than sixty jurisdictions: statutes, case law, preparatory works, each claim tied to its source. Readers of Part 1 have met this base already, as proof that understanding beats budget; it deserves its own story, because what it taught me had nothing to do with budget. The collection was done by research agents, and my instruction to them was simple: gather everything relevant to the OECD rules. They did exactly what I asked. Eighteen months of delegating to agents every day, and I still hadn't internalised the first rule of delegation: they do what you say, not what you mean. Agents even more so, and without the raised eyebrow.
The base looked rich, and I had moved on to building the interface: the part of the system that shows where a law leaves room for interpretation, with the case law that maps that room. Early one day I tested it on a jurisdiction I hadn't looked at closely, and the case-law panel came up empty. So I asked the agent the way you'd ask a colleague: you say there is case law here, and I don't see any. Did you forget to add it?
It went off and audited itself, and came back with something worse than a missed upload: the base held a reference to an Icelandic ruling with no source material behind it. We had collected the actual text of every piece of case law, so the system could recite it; here stood a citation on nothing. I asked it to go and fetch the ruling. It couldn't find it. The reference existed; the thing it referred to was nowhere the agent could see. Then I asked the question that turned out to carry this whole article: have you searched in Icelandic? It hadn't. It did, and there it was, reachable only through sources that never appear in an English search.¹
The uncomfortable part wasn't Iceland. Agents do search in local languages sometimes; ask about something Spanish and you'll often see Spanish queries go by. But they do it on impulse, not on principle, and I had assumed the impulse was enough.
If the gap existed for Icelandic, where the language is small enough to make it visible, it most likely existed everywhere.
So I gave the one instruction I hadn't known to give: go through the whole base again, and search every jurisdiction in its own language.
What one short instruction changed
The rerun started on the morning of 7 July and took the better part of a day.² There was nothing to watch while it ran, so I went back to other work. What it found, when I read it, is the clearest measurement I have made in eighteen months of this work, and the clearest confirmation that redoing everything had been the right call.
| English-led collection | Source-language collection | Change | |
|---|---|---|---|
| Raw source files | 321 | 1,783 | 5.6× |
| Case-law citations in the base | 324 | 556 | +72% |
| Preparatory works and guidance | 55 | 202 | 3.7× |
| Cells (jurisdiction × rule) | ~605 | 615 | +2% |
The number of cells barely moved, and that is the point: the structure was already there, looking finished. The rerun filled it with primary evidence. Working backwards, the English-led collection had caught roughly half of the case law and a quarter of the preparatory works.³ Fifty-seven jurisdictions gained sources. It had existed everywhere.
The examples have a pattern. Japan's court database had migrated to a search engine the agents couldn't drive, so two central rulings were reachable only through Japanese legal journals that quote the reasoning verbatim. Hungary's paragraph-by-paragraph explanatory memorandum to the law exists only in Hungarian legal databases. Poland's government rationale, a megabyte of it, which in preparatory-works terms is a novel, plus eleven administrative court rulings, live on Polish sites in Polish. And Greece delivered the strangest gift: searching in Greek proved that certain guidance doesn't exist yet, which for a legal system is itself an answer. You cannot prove absence in a language you didn't search.⁴
Whether the new material could be trusted was its own campaign: adversarial review agents went through hundreds of cells with instructions to tear the additions apart. They corrected a handful of errors per jurisdiction, and in one case the reviewer's own objection was overruled on verification. The material held, and it came back richer.⁵
Turns out it's the norm
Afterwards I went looking for whether anyone had measured this before me. The research field had, in detail. I read in English every day and had still never run into it, which is its own small proof of the problem: the papers are read by the people who build retrieval systems, not by the people who buy them or rely on them. Retrieval systems, the research shows, prefer high-resource languages and the language of the query, which in practice means English twice over.⁶ In one legal retrieval test, a standard multilingual embedding model lost half its accuracy the moment the question and the document were in different languages.⁷ The field's own benchmark for legal retrieval models is deliberately all-English. And the commercial tools are no exception: as of this summer, no vendor I could find offers genuinely cross-lingual legal retrieval; “fourteen languages” tends to mean machine translation bolted onto an English search.⁸
The blind spot has a floor under it that no model update fixes: the material itself. Sweden has published seventeen thousand decisions as guiding case law since 1981; its courts decided half a million cases last year alone. Germany publishes under one percent of its judgments.⁹ What does exist in these languages is scarce, official, and invisible from English. A model trained mostly on English text, asked in English, searching in English, will return something confident anyway. The result is not wrong answers, but a fraction of the right ones, with no signal that the rest exists.
I have since watched the same effect one layer down, in my own infrastructure. When I tested whether a knowledge index written in Swedish matched Swedish questions better than the identical content in English, the same-language version won in eight cases out of eight. A small test, and it measures direction rather than size, but the direction never flipped.¹⁰
Sometimes words are not enough
There is a limit worth being honest about, because I ran into it next. Once you search every jurisdiction in its own language, you have better material; you still can't line it up. “What is the equivalent of this Swedish provision in Germany?” is not a translation question, and text similarity answers it badly: in my blind tests on the tax base, direct embedding matching found the right counterpart 17 percent of the time. A structure that compares what the rules do, rather than what they say, found it 80 percent of the time, and on a second legal domain it did better still.¹¹ That layer is a story of its own, and a later part. The point for this one is narrower: matching language is the floor, not the ceiling.
Speak with farmers in the farmers’ language
Write the instruction down. “Search each jurisdiction, market or archive in its own language” is one line in an agent brief. The same gap sits in every market analysis: ask in English about the German market and you get the English-language slice of it. The line is not the default, and the impulse version of it is not consistent. My rerun cost a day and closed a gap I didn't know I had; the same line in the first brief would have cost nothing.
Let people work in their strongest language, and let the system do the translating. I dictate to my systems in Swedish; they query English sources in English and Polish sources in Polish. My customers' sales agent does it in reverse. The knowledge base behind it is Swedish, written by Swedish specialists; a visitor who starts writing in another language sees the whole chat switch in real time, a dedicated translation step carrying the interface and the answers while the substance stays anchored in the Swedish sources.¹² The person expresses nuance where they have it. The system crosses the border.
Build the context in, don't teach it. You can train people to add “who is the reader, which market, which language” to every request, and it helps. It survives contact with a Tuesday afternoon far better as a tool that injects that context automatically, the way a form remembers your address.
Ask your vendor the question nobody asks. None of my customers, all Swedish, buying products for a Swedish market, has ever asked whether the thing works on their language and their material. I'm not above this: it never occurred to me to ask either, until my own system failed the exam. If you operate in Europe and are being sold AI on top of an archive, ask what happens to the Hungarian contracts.
Why this is urgent
The gap compounds quietly. Every process you automate on English-led retrieval inherits the missing half without recording that it is missing; the agent reports success either way, and the coverage you didn't get never appears in any dashboard. Models will keep getting better at small languages, but the imbalance in what they were trained on, and in what exists to be found, is structural. English will stay the default. Europe's answer so far is to fund models of its own, and they may help; but no model, however European, can retrieve a judgment that was never published, and the one-line instruction is free.⁹ Iceland was just where the gap was small enough to see.
Three questions to take with you: Which of your archives, markets or obligations live in a language your AI tools never search? Who in your organisation would notice if half the sources were missing and the answers still sounded complete? What would it cost to rerun one knowledge collection with the language instruction added, and what did not adding it cost so far?
Three questions I get asked
- Isn't this just a prompt mistake? You could have written the instruction on day one.
- Yes, and that is the lesson rather than an excuse. Nobody writes the instruction they don't know is needed, and the system gives no error when it's missing; the result looks complete. The fix is knowing the failure mode exists, and after that it's one line.¹³ And if the line were obvious, one of the vendors in this summer's survey would have built it in; none had.⁸
- Won't better models make this go away?
- Partly. Models translate small languages well already, and that is the workaround. But the publication floor (a fraction of Swedish and German judgments ever becomes searchable text) sits outside the model entirely, and the English-first search habit is something a better model may soften, though you cannot see when it has. The instruction costs one line and stays useful.⁹
- Is this only a legal problem?
- No; law is just where the gap can be counted, because coverage there is countable. The mechanism is retrieval, and it applies to any archive, market or knowledge base that lives in more than one language. My two non-legal data points are small but point the same way: a Swedish-language index matched Swedish questions better in eight cases out of eight,¹⁰ and a sales agent on a Swedish knowledge base serves visitors in their own languages every day.¹²
Notes and sources
- The cell referenced an Icelandic ruling whose text and provenance were never saved; the Supreme Court and tax tribunal decisions involved turned out to be reachable only through secondary citation, the primary portals being unreadable to the agents (one renders through scripts, one blocks automated access, per the campaign's gap log). Discovered 7 July 2026.
- My own token logs, measured from session files: 2.61 billion tokens across all projects on 7 July 2026, 2.18 billion of them in this one, roughly twenty thousand novels' worth of reading and writing; a normal day in that period ran 0.7 to 1.4 billion. The collection and review campaign ran 7 to 11 July; the heaviest review day, 11 July, shows 2.55 billion in the same project. At list prices the rerun day alone prices out around 2,300 dollars, which I did not pay, thanks to flat-rate plans; a later part is about that gap.
- Measured from version control between the pre-campaign and post-campaign states: raw sources 321 to 1,783 files (22 to 146 MB); case-law citations inside the cells 324 to 556; preparatory works and official guidance 55 to 202; 580 of 605 cell files changed; 57 jurisdictions gained source-language material. The “roughly half / a quarter” figures are the share of the final citations that the English-led collection had already found (58 and 27 percent), and they overstate its coverage: 81 gap files document sources that are known to exist but remain blocked.
- All four examples from the campaign's per-jurisdiction gap logs: Japan (rulings via journals quoting reasoning verbatim, weaker chain flagged for review), Hungary (indokolás, the explanatory memorandum, via national legal databases, main portal blocked), Poland (government uzasadnienie and eleven NSA rulings via parliament and court portals), Greece (confirmed absence of substantive guidance, four Council of State rulings verified against the decisions' own numbering, and a standing rule that decision numbers cited in roundups may not be used as evidence without individual verification).
- Forty-eight review-approval commits over the campaign window, roughly ten cells per jurisdiction and about five hundred cells in all; Iceland's eleven cells yielded three factual corrections, and two citations were struck elsewhere. A later review of about 870 factual claims in a related layer found essentially all of them held: one scope error and three arithmetic slips. I treat “the reviewer got overruled once” as the healthiest data point in the set; the reviewers get reviewed too.
- Park and Lee, “Investigating Language Preference of Multilingual RAG Systems”, Findings of ACL 2025: retrievers “tend to prefer high-resource and query languages”; the double preference is exactly the trap for a European asking in English about local law.
- Amiraz et al., “The Cross-Lingual Cost: Retrieval Biases in RAG over Arabic-English Corpora”, arXiv 2507.07543 (2025): in the legal-domain test, one standard multilingual embedding model dropped from 88 to 46 percent Hit@20 when query and document languages differed; a second model degraded less, and mainly in one direction. One study, one language pair, hence the singular in the body. My own layer's language gap in comparable blind testing was about three points, for reasons that belong to a later part.
- Survey run 11 and 17 August 2026 across the major legal-AI vendors' own documentation: none ships cross-lingual functional retrieval; one's “14 languages” is machine translation over monolingual search; another's coverage tables give civil-law countries statute text only, no case law. The field benchmark (MLEB) is all-English by design.
- Sweden: 17,321 decisions published as guiding case law 1981–2026 against 520,993 cases decided in 2025 alone (Domstolsverket's annual report), about three percent of a single year's volume; commercial databases hold millions more behind licences, which is a different kind of available. Germany: under one percent of judgments published per Hamann's empirical work; the official estimate is more cautious at under three. Europe's sovereign-model initiatives address the training-data imbalance, and that matters; they do not touch the retrieval habit or this floor.
- Eight question-summary pairs, one embedding model, same proposition in Swedish and in an English translation produced by the same agent that ran the test: the same-language match scored higher in all eight, by 0.028 to 0.116 cosine, mean 0.071. Small, synthetic, one model, and it measures ranking direction rather than answer quality; I use it as a direction check, nothing more.
- Blind tests on the tax base: 1,108 statute-to-statute questions across eleven languages, answered correctly 80.3 percent of the time by the structural layer against 16.7 percent for direct embedding similarity (chance 7.9); on a second domain, data protection law across twelve jurisdictions, 98 against 24.9. The claims behind the structural layer's notes were independently adversarially reviewed: zero fabricated claims out of 1,015. How the layer works is deliberately not described here.
- The sales agent runs on a Swedish knowledge base; the chat interface is translated in real time by a dedicated translation step, so the visitor reads and writes their own language while the underlying answers come from Swedish source material.
- A later part of this series takes the general version of this lesson, what a language model actually is and why it never flags what it missed: a person cannot supply the context they don't know is missing, and the model will not ask. Language is the highest-stakes instance of it I have measured.