Recommendation Design
Edition03
AuthorFlemming Rubak
Published23 August 2026
Reading time9 minutes
On this page

The Sunday Shortlist decodes how AI makes decisions about a market category and audience.

This week we analyse UK sustainable architecture and urban masterplanning

We unpack the questions councils and developers ask AI when they choose a practice, the shortlist that comes back, and the criteria that decide it. The category has a default that is both chosen and kept.

The sharpest finding sits in the eliminator: the most severe cut in this market is insufficient expertise, in a profession that is made of nothing else. The judge is not looking at buildings. It is reading text.

Somewhere in the UK right now, a regeneration director at a local authority is asking an AI model which practice should masterplan a district. The model answers in seconds: a shortlist, the risks, a favourite. No practice hears about that conversation. We measure it.

What we measured

Since the end of March we have run this market's buying questions through the models every week: 22 weekly runs, across the decision journey from first comparison to advocacy. The article-grade window is the trailing 28 days, 27 July to 23 August. Measured with three models: Gemini, Claude, and ChatGPT.

The questions are the ones buyers ask, in their own words:

  • "What experience do you have with large-scale sustainable masterplanning projects in the UK?"
  • "How do you approach net-zero carbon design in urban regeneration schemes?"
  • "What is your track record on securing planning approval for complex mixed-use developments?"

The judging sheet

When a buyer asks a model to choose, the model behaves like a judge: it applies criteria, checks each practice's evidence against them, and cuts the practices that fail. Ten criteria decide the answers in this category. Six weigh heaviest: cost and fees, expected outcomes, expertise, product fit, regulatory and risk, and trust and reputation. The buyer language behind them is precise:

  • Expertise: "Can you evidence expertise in circular economy and regenerative design principles?"
  • Outcomes: "What measurable environmental and social outcomes can we expect?"
  • Regulatory: "How will the design comply with Building Regulations, planning policy, and environmental law?"
  • Trust: "Have you completed projects on time and within budget in the past?"

Notice the verbs: evidence, measure, demonstrate, comply. This is the exam. Every practice in the UK sits it every day, whether it knows or not.

The verdict: the funnel runs both directions

How to read the table: the percentages show how frequently each brand appears in the answers at each stage of the decision journey, measured across the full window, 27 July to 23 August. A dash means the sample for that stage is too small to report.

BrandEvaluationDecisionRetentionAdvocacy
Arup63%95%98%73%
Foster + Partners48%56%3%1%
BDP41%43%7%19%
Allies and Morrison24%7%89%87%
Shim-Sutcliffe Architects18%3%57%49%
Henning Larsen10%46%75%
Atkins8%21%27%8%

The default is Arup, and the concentration is severe: 95% of decision answers name it, and it holds through retention at 98%. The top three, Arup, Foster + Partners, and BDP, take 64.6% of all decision mentions between them. Compared, chosen, and kept: the same pattern the Danish IT market showed last week, at even higher pressure.

But this table adds something the first two editions did not have: the funnel runs both directions. Allies and Morrison appears in 7% of decision answers and 89% of retention answers, and it leads the entire category on advocacy at 87%. Henning Larsen barely registers at evaluation and owns three quarters of the advocacy answers. These are stage specialists: practices the models rarely pick, but consistently praise once the conversation turns to who clients stay with and recommend.

Last week we wrote that chosen and kept are different columns. This week goes further: every stage is its own market, with its own leaders, judged on its own evidence. A practice that has given up on cracking the shortlist still has two whole stages standing open.

One practice, two names

The window holds a finding we did not go looking for, and we state it without naming the practice. One of the category's strongest names at evaluation appears in the models' later-stage answers under two different name forms: the short form it trades under, and the longer formal one. And the two forms live different lives. At decision, the mentions split across both. At retention, one form carries 83% of the answers while the other carries 15%.

83% and 15%

One practice's retention-stage presence, split across the two name forms the models learned it under. Same practice. Same buildings. Two entries in the ledger.

We saw a version of this last week, when consolidated players appeared in the answers only through the brands they had acquired. This is the same mechanism one level down: the models' recommendation ledger is keyed to name strings, not to firms.

The mechanism is worth spelling out, because it decides the fix. A model has no registry of companies. It has text. It learned your brand from every page, article, register entry, and award citation where your name appeared, and the exact string on those pages is the handle the learning attaches to. Retrieval and citation follow the same rule: a citation binds to the words actually on the page.2 If half of your coverage carries one name form and half carries another, the capital splits. Neither entry carries your full weight into the answer, and no dashboard will ever show you the split.

The exposure class is far bigger than one practice. Renamed firms, merged firms, acquired brands, abbreviations, trading names versus legal names, the "& Partners" and "Group" and "Studio" variants: every one of them is a potential second entry in the ledger. Brands maintain style guides for how humans should write their name. Almost none have ever checked which form the models keep their reputation under.

Consolidating is unglamorous and cheap, and it looks like this: pick one canonical form and use it everywhere text is written, from bylines and award submissions to press boilerplate and the professional registers. Declare the equivalence in machine-readable form, Organization schema with alternateName and sameAs on your own domain, so parsers can join what the ledger holds apart. Then audit it: ask each model about every name you have ever traded under, and read which entry the answer draws on.

Check which name AI actually recommends before you assume the ledger is whole.

Nobody owns a criterion

Ten criteria, zero owners. No practice is the default answer for measurable outcomes, for regulatory compliance, for circular-economy expertise, for transparent fees. The default wins on coverage and reputation, not on owning the questions.

That is now three markets in a row: US CRM platforms, Danish IT outsourcing, UK sustainable architecture. Three categories on three different markets, software to services, and not one of the thirty measured criteria has a default answer attached to a brand. The vacuum is starting to look structural, and it is the standing opening in every category we decode: the first brand that publishes real evidence against a single criterion takes it nearly uncontested.

The eliminator: insufficient expertise

The most severe elimination trigger in this category is not cost. All three models cut practices on insufficient expertise, and the buyer phrase behind it is blunt:

"Lack of specific experience in sustainable design principles for complex urban environments is a critical gap."

Stop on that for a moment. This is a profession of chartered experts, decades-long portfolios, and built proof standing in public view. And the judge's most severe cut, in the aggregate, found the expertise evidence lacking in five of six candidate profiles in this run.

The explanation is the finding of the week. The judge has never seen a building. The models cannot walk a completed district, photograph a facade, or feel a public realm work. They read text, and they cite text. A practice's expertise lives in its buildings, its drawings, its photography, its awards shelf. Almost none of that is text a model can lift as evidence against the question "can you evidence expertise in circular economy and regenerative design principles?"

An architecture portfolio is an argument made of images, presented to a judge that reads. The practices in this category publish pictures of their proof and captions written for humans who can see. The model meets a beautiful page and finds nothing it can quote. The expertise is real. The evidence, in the only format the judge accepts, is missing. That is how a category made of expertise gets cut for insufficient expertise.

What claiming it looks like

This is the working brief for that page, generated from the monitoring data, with the brand genericised. The frame is "show the evidence": this eliminator is proof-shaped, so the page opens with the hardest certified proof point and lets the methodology explain why the outcome repeats.

Title: "Does [my brand] Lack Urban Sustainability Experience? The Evidence"

The position the page takes: [my brand] is one of the category's most evidence-rich practices for sustainable urban design, with measurable performance outcomes across complex city-scale projects. Defensible because it rests on named projects with third-party-verified certifications and an in-house environmental analysis capability, not on portfolio adjectives.

The intent family the same page must also answer (one prompt spawns two to three searches behind the scenes):

  • "What sustainable urban projects has [my brand] completed?"
  • "How does [my brand] approach environmental performance in large-scale urban design?"
  • "Which architecture firms have proven experience in sustainable masterplanning for cities?"

The structure: five H2s, each a claim the models can lift as a standalone answer:

  • "The objection stated plainly: where the 'experience gap' concern comes from and why it does not hold"
  • "[My brand]'s environmental analysis capability: how sustainability is embedded at concept stage, not retrofitted"
  • "Named projects, verified outcomes: the schemes that prove the methodology works"
  • "How [my brand]'s approach differs from firms that treat certification as the endpoint"
  • "What to ask any architecture firm to verify sustainable urban design capability, and how [my brand] answers each question"

Key Takeaways for the top of the page (each a self-contained claim a model can cite):

  • "Decades of urban sustainability delivery: environmental performance has been integrated into complex urban projects since the practice's early masterplans. Not adjacent experience: the core of the practice."
  • "In-house environmental analysis: climate modelling, energy simulation, and carbon assessment run from the earliest design stage, structurally different from firms that outsource sustainability to consultants."
  • "Verified certifications on named projects: certification scores and zero-carbon frameworks that are third-party verified outcomes, not claims."
  • "Urban complexity at scale: sustainable delivery across extreme climates, high-density contexts, and politically complex environments. Complexity as the standard operating condition."

One note on the last H2: handing buyers the verification checklist every rival must also answer is the generous move that only the practice with the evidence can afford to make.

Key Snippet, placed early: "[My brand] has led sustainable urban design across [N] countries, with verified net-zero outcomes on major city-scale projects."

Slug: /sustainable-urban-design-experience-complex-environments — the buyer's question shape, not a content-type label.

Kept alive: certifications re-verified against the issuing bodies, the latest sustainability report cited within 24 months, the refresh dated on the page.

Where the evidence must live: who listens where

The models name their sources, and this week's list reads like a public-sector procurement file. Across the three models, the answers cited ARB (the Architects Registration Board), RIBA, the RIBA Journal, Crown Commercial Service, the Construction Industry Council, Hansard, and the National Audit Office. The regulator, the professional bodies, the procurement framework, the trade press, and Parliament's own record. Not one practice website.

That list is not random, and the research explains the mechanism behind each move you should make:

1. Your evidence must exist as text, and it must parse

Network-traffic captures show the models going to the official page first for factual claims and giving up when content hides behind JavaScript; in one recorded reasoning trace, ChatGPT wanted a vendor's own numbers, could not parse the page, and cited a third-party source instead.2

For a practice, this is the portfolio problem stated as engineering. A project page that is a photo essay with a lyrical paragraph gives the judge nothing. The same project written as evidence does: named project, completion date, measured outcomes (energy performance against target, carbon figures, planning consent secured, post-occupancy results), the team's accreditations, in plain HTML.

One page per claim, because search results dedupe by domain: twenty thin project pages collapse into one candidate, while one strong evidence page stands alone.2

2. The verdict about you needs third-party surfaces

Vendor pages get cited for their own facts; the judgment gets cited to third parties.2 This week the third-party surfaces are named in our data: the RIBA Journal and the professional bodies carry the editorial layer, and awards with written citations (not logos on a carousel) put your expertise in someone else's text.

3. Match the surface to the engine

The per-engine differences are measured.

ChatGPT and Claude lean on text surfaces: professional-body pages, trade press, LinkedIn articles under named authors. Video is close to citation-dead in ChatGPT, because search fetches a video's metadata, not its transcript.2,5

Gemini and Perplexity read the spoken word, transcribed. Gemini reads Google's surfaces, YouTube transcripts included; Perplexity quotes video heavily and rewards fresh, dated material.4 A practice sits on a natural asset here: the project walkthrough. One take of a partner walking a completed scheme, explaining the sustainability strategy in the buyer's vocabulary, transcribable, is the portfolio translated into the judge's format.

4. The registers are evidence you already own

The models cited ARB and Crown Commercial Service this week. Your ARB registrations, RIBA chartered status, framework appointments, and filed accounts are machine-readable trust signals that cost nothing to keep clean and current. Check what the registers say about you before the models do.

5. Date everything

In an analysis of 250 million AI responses across eight answer engines, half of all top-cited content was under 13 weeks old.6 The recency bias is structural, and the researchers reading the engines' network traffic reached the same standing rule from the other direction: date every claim.3

Post-occupancy data is the architecture category's unfair advantage here: a completed building keeps producing fresh, dated, measurable evidence for decades, if anyone writes it down. A project page whose performance figures carry this year's date beats a timeless one twice: once with the buyer, once with the judge.

The lesson

In this category the expertise is real, the buildings stand, and the judge still cuts on insufficient expertise, because the evidence lives in a format the judge cannot read. The fight is not to become more expert. The fight is to translate expertise you already have into text a model can quote, against criteria nobody owns, in a market where the consideration shortlist has not moved in eight weekly runs.1

Your category has its own version of this finding. The unreadable evidence differs. The translation job does not.

Sources

The network-traffic findings below come from one researcher's logged-in accounts and are labelled directional by the author; the mechanisms are reproducible, the percentages are not population measurements. Capture dates matter: the plumbing changes faster than the mechanisms.

  1. Suganthan Mohanadasan, ChatGPT Already Knows Who It'll Recommend Before It Searches, suganthan.com, August 2026.
  2. Suganthan Mohanadasan, How ChatGPT Actually Picks Sources (I Read the Network Traffic, Not the Outputs), suganthan.com, June 2026.
  3. Suganthan Mohanadasan, ChatGPT Changed How It Picks Sources While You Were Reading My Last Post, suganthan.com, July 2026.
  4. Suganthan Mohanadasan, How Perplexity Actually Picks Sources (I Read the Stream, Not the Answers), suganthan.com, July 2026.
  5. Ahrefs, Why ChatGPT Cites Pages, analysis of 1.4 million ChatGPT prompts, 2026.
  6. Josh Blyskal, Profound, We Analyzed 250 Million AI Search Results, analysis across eight answer engines, 2025.

The Sunday Shortlist

Get the decode, and the toolkit

The Sunday Shortlist decodes one category at a time: the questions buyers ask AI, the shortlist that comes back, and the criteria that decide it. Subscribers get every decode first, and the free toolkit gets you started tonight: a worksheet and three prompts to run the first check on your own brand.

One email a week · unsubscribe anytime · Terms

All editions