Lenswork: perspective matters.
Ask a language model on its own and you get one perspective, the middle, and it reads as finished. Over eighteen months of daily work with these models, I've ended up with a way of bringing the other perspectives back in. It grew out of the work, not out of a plan. Then I tested it. It holds up, as long as someone weighs what comes back.

One model, one perspective
I don't trust a first draft from a model. Not because it's bad. Because it's plausible, and plausible is the hardest thing to catch.
A language model predicts the most likely continuation of what you gave it. Ask it a question without saying where you stand, and you get the most likely answer to that question: the general one. It comes back well structured, confident and complete. Nothing in it tells you what is missing, because an answer from the middle has no visible holes.
That is fine for a lot of work. If the question is simple, the middle is usually where the answer is. But much of what people now do with these models is knowledge work: a strategy, an analysis, a recommendation, a decision paper. There the middle is a problem. In Part 5 I gave eleven models from three vendors the same open go-to-market question. All eleven picked the same target group, the same positioning and the same price band. They even found the same gap in the market.¹ If everyone with the same tool finds the same gap, it isn't a gap. It's a queue.
Part of this is how the models are made. The tuning that makes them helpful and polite also narrows them: in one controlled comparison, the spread of likely next words fell by about a third after that tuning, and the outputs became about twice as similar to each other.² Some of that narrowing is deliberate, and much of it is good. Popular opinion isn't truth. But it means the default answer is built to converge, and if you work on questions where the value is in being different, converging is the one thing you don't want.
Think of it as white light. Everything the model knows is in there, but shine it straight ahead and you get one even glow. The colours only show when something splits it, the way a prism does. That is what the lenses are for.
A field experiment at Procter & Gamble says the same thing from the other side. People working alone with AI did as well as teams working without it.³ It's easy to read that as the model replacing the team. I read it differently. What a team gives you is the other perspectives, and one person with a model can get them back, but only by asking for them. Left alone, the model gives you the middle.
None of this makes the model useless for strategy. It's very good at structuring, drafting, summarising and keeping track of what you've decided. What it doesn't do is the weighing. The decision is still yours. The question is how to give yourself more to weigh.
How i work now
I call it the Lenswork Method. There are three steps, and the order matters.
First, you make a finished draft. With a model or without one, it doesn't matter. What matters is that you've taken it as far as you can on your own: you know what you think, you've brought in the background and the constraints, and you'd be willing to show it to someone. The better the draft, the better everything that comes after it. A lens can't sharpen a thought that isn't there yet. Ethan Mollick gives the same advice from another direction: ask the AI for feedback when the draft is done, not before.⁴ How to get to a good draft with a model is a subject of its own, and I won't go into it here. On the articles in this series, that phase alone is usually seven to fifteen rounds.
Second, when you're done, you read it through lenses. A lens is a role, not a person. Not "Karl, 38, lives in Lisbon, cycles at weekends, works as a tech lead at a food company". A CFO. The person in procurement who will actually buy the thing. A tax specialist. The process owner whose day changes if this goes through. You ask the model to read your draft the way someone in that role would, with that role's knowledge and worries, and to say what it would question, what it would need, and what is missing. Run each lens in its own, fresh conversation, so one reading doesn't colour the next.
The difference isn't cosmetic. A model can only give back what it was trained on. There are forums, trade journals, networks and decades of writing about how a CFO looks at an investment or how procurement reads a proposal. There is almost nothing about Karl. Describe a role and the model has a whole profession to draw from. Describe a person and it has to invent one. That's my reading of why roles work and personas mostly don't, and it's also the limit of what a lens does: it gives you a perspective, not more knowledge. A model playing a CFO doesn't know more facts than the same model playing nobody.⁵ It looks at the same facts from somewhere else.
The lenses aren't a fixed set. You choose them for the question in front of you: who will receive this, who is affected by it, which markets or segments it touches, which functions have to say yes. When I built a product page for a tax reporting tool, I read it through a CFO, an in-house tax manager, procurement, and a management consultant who might use the tool for clients. They wanted different things from the same page. The page got better in places I wouldn't have looked.
Most of the time a lens changes details: a sentence that a CFO would read as a commitment, a number procurement will ask for, a risk you mentioned once and should have mentioned twice. Occasionally it changes the whole thing. That's rarer than you'd think, and it's worth the rest.
A lens can be wrong. So can a colleague. The point of a team was never that everyone in it is right. It's that each of them makes you look at the problem from somewhere you weren't standing.
Third, if the question is strategic, long-lived or critical, you take the finished draft to a model from another vendor. In a way it's one more lens: same question, different training, different habits. At some point you and your model reach the end: it tells you, in effect, that the draft is done and there is nothing more to get. That's when a second model earns its keep. But only if it knows what you know. Ask the first model to write a context file: the brief, the goals, the key facts and constraints, who will receive the document and who is affected by it, and the trade-offs you made on purpose, including feedback you chose not to follow. Send the draft and the context file together. You rarely need to name the client to do this: describe instead of naming, and remember that what is publicly known about a company isn't customer data. On a business agreement through the API, the vendor doesn't train on what you send. In my experience, send the draft alone and you get a confident review of a document the reviewer doesn't understand.
Also in my experience: somewhere between half and two thirds of what the second model raises is worth taking. Not all of it. That's fine. You're the one deciding.
The templates, and the three cases from my tests as worked examples (context, lenses and what came out), are in the Lenswork Method toolkit.
When it’s worth it
Not every time.
Use it when something matters: a bigger decision, a strategy that has to hold for years, a document that goes out under your name, a moment where getting it wrong is expensive. That's where the step is easy to skip and costly to miss, because it shows you things you wouldn't have seen otherwise.
For a simple question, ask the model once. Simple questions don't have much to untangle, and as the test below shows, models agree with each other more the simpler the question gets. Lenses cost time and tokens, and a second model costs more. This isn't a framework for every prompt you write. It's for the ones you'd otherwise have taken to a meeting.
Does it actually do anything?
It's easy to believe that asking from different angles gives you different answers. It's also easy to believe you're just getting the same answer in different clothes. So I measured it.
Three tasks, from simple to complex. A simple one: should you follow up with a client who hasn't replied a week after a sales meeting. A middling one: should a company use a particular brand name. And a complex one: write a legal email in a dispute. Six models from two vendors, Claude and Gemini. Each task asked with no lens and through several professional lenses, repeated so that each model's own variation could be measured, and run in isolation with nothing else in the model's context. The answers were coded blind: whoever coded them never knew which model or lens had written what.
What the test shows is how the tool behaves, not which answer is right. Two things.
Lenses move the answer. On the legal email, within every one of the six models, answers from different lenses were 1.7 to 2.2 times further apart than three runs of the same lens. That second number matters: a model never gives you exactly the same answer twice, so without measuring how much it varies on its own, a lens effect means nothing. The lenses beat that noise in all six. The two lawyer lenses turned a complaint into a formal demand; the lens built from a litigator and a mediator turned the same email towards keeping the relationship. What the sender actually asked for barely changed. The lens changed how he asked.
And the harder the question, the more models disagree. On the simplest one, 59 of 60 answers said yes, follow up, and the models differed from each other only about a quarter more than each differed from itself. Even the lenses barely moved it. On the brand name, about 40 to 55 per cent more, and nearly all of them still landed on the same advice. On the legal email, 85 to 90 per cent more, with or without a lens, and the vendors split: one family negotiated, the other refused.
That is the case for step three. On a simple question a second model tells you what the first one said. On a hard one, it may not.
Three tasks are not a law. The first two were measured the same way, the legal email with a different instrument.⁶ But the direction holds at every step, and it's what I see in my own work.
What happens without you
Then I took myself out of it. Three made-up but realistic tasks: a proposal to win a procurement, a board paper on changing a company's pricing model, and a company's statements after a food recall in which two people had died. Four models each got the brief and wrote a first draft. Seven or eight lenses per task read it, the model revised, the lenses read it again, the model revised again, wrote a context file, and sent it to a model from the other vendor. The author decided whether to take the feedback. At the end a third model, outside the chain, compared the first and the final draft without knowing which was which.
The final draft won nine of twelve outright, and one more was a split decision.⁷ That is a useful result on its own: run the method automatically and you will usually start your own work from a better place than the first draft.
The two losses are more interesting. In the recall, the lawyer and prosecutor lenses pulled the statement towards caution, the affected families and the journalist pulled it towards openness, and the model, weighing alone, chose caution. The final version mentioned "an environmental sample result from early September" and declined to discuss timelines. The first draft had said plainly that a positive test was missed for a week. The judge preferred the first. In the procurement, the draft absorbed so much feedback that it started promising things the seller couldn't deliver.
It went the other way too. In the procurement chain run by Gemini 3.1 Pro, the first draft had the buyer's truck fleet wrong. Fourteen lens reviews read past it. The second model, from the other vendor, caught it in its first paragraph.
Many drafts also grew. Every lens added something, nothing took anything away, and with no one holding the word limit, some came back three times the length they were asked for. With a person in the loop, that's a five-second fix. Without one, nobody makes it.
I also tried the obvious improvement. People in a team get further because they talk to each other, so maybe the lenses should too. I ran the same tasks three ways: lenses reviewing on their own, lenses reading each other's reviews and agreeing on a joint list, and a chair summarising for them. With three of the four models, blind judges preferred the lenses that worked alone. With the fourth, the discussion came out ahead. Every time, it cost about six times as many tokens.⁸ Whatever a team of people gets from talking doesn't carry over. With models, the perspectives come from naming the lenses. Letting them talk only added cost.
So the method gets a model further than its first draft, often a good deal further. But where the lenses pull in different directions, it can also fold in on the wrong side, because nobody is there to choose. That is the job that stays with you. The lenses give you the tension. You decide which side of it you're on, and what to leave out. The strongest result isn't the model on its own, and it isn't you on your own. It's the two of you, each making the other better.
On monday
Make a finished draft, with or without a model, as far as you can take it yourself.
Then pick two to four lenses: the people who will receive it, decide on it or be affected by it. Ask the model to read the draft as each of them, one fresh conversation per lens. Take what makes it better. Leave what doesn't.
If the question is complex or critical, ask the model for a context file and send it, with the draft, to a model from another vendor. See what's different.
Then decide. That part was always yours.
If you want it step by step, with the templates and the examples from the tests, the toolkit is here.
Notes and sources
- Part 5 of this series. One open go-to-market question, eleven models from three vendors over four generations, no format requirements.
- Mohammadi (2024), Llama-2 7B base against Llama-2 7B chat: same architecture and pretraining data, only the tuning differs. Entropy over the top five next tokens fell 35 per cent; cosine similarity between outputs rose 2.2 times. One model pair; the size of the effect varies between models.
- Dell'Acqua, Ayoubi, Lakhani, Lifshitz, Sadun, L. Mollick, E. Mollick et al. (2025), "The Cybernetic Teammate", a field experiment with professionals at Procter & Gamble, summarised by Ethan Mollick in One Useful Thing (22 March 2025). Without AI, teams beat individuals by 0.24 standard deviations; individuals with AI improved by 0.37, as well as the teams. The study doesn't test lenses.
- Ethan Mollick, One Useful Thing (oneusefulthing.org/about): he asks for AI feedback only after finishing a complete draft.
- Zheng et al. (2023): 162 roles, many of them occupations, tested on 2,410 factual questions, with no reliable improvement. Basil, Mollick et al., "Playing Pretend: Expert Personas Don't Improve Factual Accuracy" (Wharton, 2025): "you are a physics expert" did not help on graduate-level physics questions. Both test whether a role makes the model more correct. Neither tests whether it changes what the model looks at, which is what a lens is for.
- Legal task: 72 emails, six models, four lens conditions, three runs each, coded blind by six coders on six dimensions, with an independent second coder on twelve (agreement 0.06 on a 0–1 distance scale against 0.17–0.25 between identical runs). Brand-name task: 60 answers, coded blind against a concept ontology. Follow-up task: 60 answers, same design and instrument as the brand-name task, with the ontology also built blind. Ratios are distance between models divided by distance within a model. How the tests were run is shown in the worked examples in the toolkit, so you can run them yourself.
- Three scenarios, four authoring models (Opus 5.5, Fable 5.1, Gemini 3.8 Flash, Gemini 3.1 Pro), one run each. Judges outside the chain, blind to which draft was which: Gemini 3.8 Flash for the Claude chains (both orders; in one of the six, the recall chain written by Opus 5.5, the verdict followed whichever draft it read first), Fable 5.1 for the Gemini chains. The two judges weigh length differently, so levels between vendors aren't comparable; the direction within each chain is. Every step is saved.
- Three scenarios, four authoring models, two runs each, one lens round, blind pairwise judgements by two judges outside each chain (for the Claude chains Fable or Opus plus Gemini 3.8 Flash). Lenses alone beat the discussion 63 to 33 overall; the discussion won only with Fable 5.1 as author. Points first raised by a single lens survived equally often whichever way the lenses worked; what decided whether they reached the final draft was the authoring model. Tokens per condition: 1.8 million alone, 11.1 million with discussion.