The source ecosystem: how AI actually decides who to name

When we asked five assistants who to hire in our own category, 59 of 108 citations were other people's roundup lists. Five of the twelve most-named companies sat on one directory. Being named is mostly about other people's pages.

Tom HardimanUpdated 17 August 20263 min read

Almost all AI visibility advice is about your own website. Fix your headings, add schema, publish more pages. Our own measurement says the majority of what decides whether an assistant names you is not on your website at all.

What we measured

We ran our own business through our own audit tool: three questions our buyers actually ask, across ChatGPT, Claude, Gemini, Google AI Overviews and Perplexity. 29 usable answers. We recorded every source URL, then categorised them.

Type of page citedCountShare of 287 citation URLs
Blog posts10938%
How-to guides and tutorials6422%
Best-of and roundup lists5419%
Forum and community threads62%
Explainers and definitions21%
Case studies00%
Grade B — credible testing, methodology shown

Not one case study was cited, in 287 citations. The content type agencies most want to produce, and most often sell as proof, was used zero times by any assistant answering any of the three questions.

The vendor question runs on other people's lists

Split by question, the pattern sharpens. Two of our three questions were how-to questions, and the assistants answered them with advice and named no companies at all. Only the explicit 'who can I hire' question produced vendors.

On that question, 59 of 108 citations were third-party roundups — pages titled things like 'top 10 GEO agencies', 'best AEO agencies UK', or a directory listing. The assistants were not evaluating agencies. They were reading somebody else's list of agencies and repeating it.

Grade B — credible testing, methodology shown

Five of the twelve companies named two or more times appeared on a single agency directory that the assistants cited. The company named most often appeared on all three of the roundup pages we could access. We appeared on none of them, and were named zero times.

That is the mechanism, visible in one run. If you want to be named on a hiring question, the work is largely getting onto the pages that already answer it.

Which sources carried the most weight

SourceAnswers it fed (of 29)Assistants citing it
reddit.com6Gemini, AI Overviews, Perplexity
semrush.com4Gemini, Perplexity
eseospace.com4Claude, Gemini, AI Overviews
impactplus.com4Claude, Gemini, Perplexity
linkedin.com4AI Overviews, Perplexity
entrepreneur.com3Claude, AI Overviews, Perplexity

Reddit came first, which matches every large-scale study of this. Reddit is the single most-cited domain across the major engines, and Perplexity in particular leans on it heavily — reported at around 46.7% of its top-ten citation share.

Reddit being the most-read source is not an invitation to astroturf it. Posting promotional material disguised as a personal recommendation breaks Reddit's rules and creates a reputational problem that outlasts any visibility gain. The usable version is answering questions honestly, with your affiliation stated.

Why off-site beats on-site

The independent evidence points the same way as our run. Ahrefs measured 75,000 brands and found unlinked mentions of a brand elsewhere on the web correlate with AI visibility at r=0.664, against r=0.218 for total backlinks — roughly three times stronger.

Both figures are correlational and the direction of causation is genuinely unclear, because large brands get mentioned and get cited. But the ordering is consistent: what other people's pages say about you outweighs what your own pages say about you.

What to actually do

  1. Find the roundups for your category. Run the questions your buyers ask, record which pages the assistants cite, and list every roundup among them. That list is your target set, and it is specific to your category rather than generic.
  2. Get listed where listing is possible. Some cited sources are directories that accept submissions. That is the cheapest available action and it requires nobody's permission.
  3. Pitch the editorial ones. Roundups written by other agencies or publishers are a pitch, not a form. They are still worth the email.
  4. Answer questions where your buyers ask them, under your own name. Reddit is first in the data for a reason, and honesty is the only sustainable way to be there.
  5. Keep your listing details consistent. For local businesses especially, inconsistent name, address and phone details across directories directly degrade what assistants say about you.

The thing this replaces is the instinct to publish more pages on your own site and wait. On the evidence here, that is the slower half of the job.

Limits

This is one run, in one category, on one day: 29 answers and 287 citation URLs. The zero case studies and the 59-of-108 roundup share are strong signals, but a single run cannot tell you the precise proportions, and a different category would produce a different source mix. Treat the method as the transferable part, not the percentages.

Common questions

What is the source ecosystem in AI search?
It is the set of third-party pages an assistant reads before answering — directories, roundups, forums, publishers and review sites. In our own measured run, the majority of what determined whether a company was named sat on those pages rather than on the company's own website.
Do case studies help you get cited by AI?
In our run, no. Across 287 citation URLs from five assistants, exactly zero were case studies. Blog posts, how-to guides and roundup lists made up 79% of everything cited.
How do I get on the lists that AI assistants read?
Identify them first by running your buyers' questions and recording which pages get cited. Some are directories that accept submissions, which is the cheapest route. The rest are editorial roundups that require a pitch to whoever wrote them.

Sources

Every claim above should be checkable. Where a study has limits, they are stated rather than left out.

  1. Sentinay — original measurement, 29 answers and 287 citation URLs across five assistants, 17 August 2026Grade B. Our own single run; limits stated in the guide.
  2. Ahrefs — 75,000 brands, unlinked mentions r=0.664 versus backlinks r=0.218Grade B. Correlational.
  3. Profound — citation share analysis across 680 million citationsGrade B. Large-sample observational.

More on getting cited in ai answers

  • Why the five AI assistants disagree about who to recommend

    We measured every source five assistants used to answer the same questions. Not one domain was cited by all five, and 89% were cited by only one. Our own original data on why a single AI visibility number is misleading.

  • What actually works in AI visibility, graded by evidence

    Twenty interventions sold as AI visibility work, each rated by how strong the evidence behind it actually is. Three are gates you must pass. One has a rigorous causal study behind it, and that study found nothing.

  • How AI assistants pick which businesses to name

    ChatGPT, Perplexity, Google's AI answers and Copilot use four different pipelines to decide who gets recommended. An intervention that works on one can do nothing on another, which is why they have to be measured separately.

  • What to fix first for AI search

    A running order based on what the evidence supports rather than what is easiest to sell. Three gates before anything else, then off-site work, then your own pages — which come later than almost every agency will tell you.

  • Does llms.txt do anything?

    Over 844,000 sites have adopted llms.txt. No major AI provider has committed to reading it in production retrieval, and the largest correlation study found no relationship with citations. Here is what it is actually for.

Want to know where you stand?

We ask five AI assistants for a company like yours and send you a free report showing how often you came up and who came up instead.

Get my free report