TL;DR

  • Three AI search engines, one honest answer: none of them documents how it decides which eligible page to cite. Being reachable by their crawler is documented. Being chosen is not.
  • Google admits the most. It says appearing in AI Overviews and AI Mode is “still SEO,” and it publishes a list of things you do not need: no llms.txt, no special schema, no content chunking, no AI-specific writing style. OpenAI and Perplexity document their crawler and stop.
  • Every enrichment tactic sold as “GEO” is a cross-engine hedge, not a proven lever. The one paper that measured them reported gains up to 40 percent, but those are its own benchmark numbers, and on Google’s surface Google explicitly says skip the markup theater.
  • The tell of a competent partner: they label the hedges as hedges. Anyone selling a guaranteed citation in ChatGPT is selling a mechanism no vendor has ever published.
  • Apply the enrichment where selection is undocumented (ChatGPT, Perplexity). Drop the markup theater where Google tells you it is ignored. Never price either as a guarantee.

There is a layer of AI search that every vendor documents in detail, and a layer none of them documents at all. Most of the overselling in this market comes from treating the second as if it were the first, so it is worth pinning down which is which, with each vendor quoted against itself.

AI search has three layers. Eligibility (can the engine fetch you), citation selection (does it choose you), and parametric memory (does it already know you without looking). We covered the whole model in our GEO field notes. The middle layer is the one everyone gets wrong, so that is where this sits.

What is citation selection, and why can nobody document it?

Citation selection is the step where an AI answer, having gathered a set of pages it is allowed to use, decides which of them to quote and link. Eligibility is the layer beneath it: each engine only considers content its own named crawler can fetch, and that layer is fully first-party. Google names Googlebot, OpenAI names OAI-SearchBot, Perplexity names PerplexityBot. Confirm those are allowed, in robots.txt and at your firewall, and you have cleared the documented part. We walk through that check in the AI SEO mandate.

The selection step is different. No vendor publishes how the model ranks and picks among the eligible sources. Not Google, not OpenAI, not Perplexity. Everything you have read about “optimizing for AI citations” is an inference about this layer, drawn from research and observation, not from a specification any of the three has released. None of which discredits the research. It just means holding it at the confidence it has earned, and staying suspicious of anyone who states it as fact.

What does Google actually admit?

Google goes furthest of the three, and its position is deflationary on purpose. In its AI-optimization guidance, Google frames AI Overviews and AI Mode as retrieval over its ordinary Search index plus query fan-out. In plain terms: it is still SEO. A page is eligible only if it is indexed and snippet-eligible, the same bar as classic Search.

The more useful part is what Google says you do not need. It lists tactics the AI-SEO cottage industry sells, and tells you to skip them on its surfaces:

  • An llms.txt file, special markup, or Markdown for AI. In Google’s words, “Google Search ignores them.”
  • Special structured data. Google states that schema “isn’t required for generative AI search, and there’s no special schema.org markup you need to add.”
  • Chunking your content into tiny AI-sized pieces.
  • An AI-specific writing style.
  • Inauthentic “mentions” seeded across the web to look talked-about.

Read that list next to any GEO pitch deck. Several of the line items being sold as AI-citation levers are things the largest AI-search surface on earth tells you, in its own documentation, that it ignores. Worth carrying into your next vendor conversation.

One caveat, and it is Google’s own: this list is scoped to Google’s surfaces. It says nothing about ChatGPT or Perplexity, and neither of those vendors has published a matching “not needed” list. Do not let a “Google says schema is optional” become “schema is useless everywhere.” That overreach is its own kind of dishonesty.

What do OpenAI and Perplexity leave silent?

Both document their plumbing cleanly, then stop at the interesting question. OpenAI splits its crawlers into three, and only one governs whether ChatGPT can cite you:

  • OAI-SearchBot is the search and citation crawler. Opt it out and your site will not appear in ChatGPT’s search answers.
  • GPTBot is training only. Blocking it means “do not train on my content,” and has no effect on whether you can be cited.
  • ChatGPT-User is a user-triggered fetch mid-conversation, not systematic crawling, and it does not determine search eligibility.

Perplexity mirrors the pattern. PerplexityBot is the crawler that surfaces and links you in results, and it is explicitly not used for training. Perplexity-User is the user-triggered fetch. Both vendors also publish their crawler IP ranges (at openai.com/searchbot.json and perplexity.com/perplexitybot.json), because allowing the user agent in robots.txt is not enough if your Cloudflare or AWS firewall is quietly blocking the actual traffic. That edge-layer gotcha sinks more sites than robots.txt ever does.

Notice what is present and what is absent. Present: exactly which bot controls eligibility, and how to let it in. Absent: any statement of how, among the pages it is allowed to read, the model decides whom to cite. They document the door and say nothing about the selection behind it.

The honest confidence map

The whole thing fits on one grid. Watch the middle column: on every surface, selection is inferred, never documented. The gap is not in your knowledge, it is in what exists to be known.

SurfaceEligibility gate (first-party, documented)How it picks citationsWhat the vendor tells you to skip
Google AI Overviews / AI ModeGooglebot indexing, page must be snippet-eligibleNot documented. Google only says it is “still SEO”llms.txt, special schema, chunking, AI writing style, seeded mentions
ChatGPT searchOAI-SearchBot allowed in robots.txt and at the firewallNot documentedNothing published
PerplexityPerplexityBot allowed in robots.txt and at the firewallNot documentedNothing published

Any agency can fill in the first two columns from public docs in an afternoon. The third column is where competence shows, because it takes reading the primary sources, and the middle column is where honesty shows, because the temptation is to pretend you have filled it in.

So what do you actually do about the layer nobody documents?

You hedge, and you say out loud that you are hedging. The most-cited research on this is the original GEO paper (Princeton and IIT Delhi, 2023). On the authors’ own benchmark, targeted content edits raised visibility inside generative answers by up to 40 percent, with the biggest levers being:

  • Adding relevant quotations from credible sources: +27.8 percent.
  • Backing claims with concrete statistics: +25.9 percent.
  • Cleaner, more readable prose: +25.1 percent.
  • Attributing claims to cited sources: +24.9 percent, and up to +115 percent for lower-ranked sites, which is why this layer is a leveler for smaller brands.

Those numbers are worth acting on. They are also one paper’s reported results on one benchmark, not a vendor specification and not independently reproduced here, so they belong in the “reasonable hedge” column, not the “proven” one. And they collide directly with Google’s guidance: the same enrichment that plausibly helps on ChatGPT and Perplexity is, on Google’s surface, partly the markup theater Google tells you to skip. A tactic that helps on one engine and is ignored on another is an engine-specific bet, not a law of AI search.

So the working rule is narrow. Apply the enrichment (real quotes, real statistics, clean attribution, readable prose) where selection is undocumented and plausibly rewards it: ChatGPT and Perplexity. Do not over-build markup for Google, which says it is ignored. And never convert any of it into a promise, because none of it rests on a mechanism a vendor has published.

Why on-page work has a ceiling

There is a further reason to keep the enrichment tactics in proportion. A 2026 study of 167,551 AI citations across 128 brands and 12 markets found that 85.7 percent of citations point to third-party sites the brand does not own, and only 14.3 percent to owned pages. Roughly 80 percent of citations came from about 18 percent of domains, and Wikipedia was the most-cited domain in 11 of 12 languages.

That reframes the exercise. Polishing your own pages is necessary and worth doing, but the citations themselves are overwhelmingly earned somewhere else, on sources the model already trusts. Which points at the third layer, the model’s trained-in memory, where no crawler setting or on-page hedge reaches at all. We take that layer apart in the next post in this series. For now, keep the proportion in view: on-page GEO is the part you control, and the smaller half of what decides the outcome.

How to buy AI-search help without getting sold a mechanism

You do not need to become an expert in any of this to hire well. You need one test. Ask the agency how AI engines decide which source to cite. A competent answer separates the documented layer from the inferred one, quotes Google’s “still SEO” position, names the crawlers, and labels the enrichment tactics as hedges with a confidence level attached. An incompetent or dishonest answer states the selection mechanism as fact, or worse, guarantees you a citation in ChatGPT.

No one can guarantee that, and the honest firms will tell you so before you ask. This is the same posture we bring to a GEO engagement and to every audit we run: fix the documented layer completely, hedge the inferred layer deliberately, and price the difference between the two honestly. If you want to talk through where your own site sits on this map, that is what a consultant engagement is for.

Plenty of firms have collapsed all three layers into one confident pitch. The ones who know what they are doing will tell you which part they cannot promise, usually before you think to ask.