TL;DR
- A new preprint strips the two easy explanations out of the citation decision. Every candidate was real, equally visible, and stripped of status signals. Eleven models from three vendors still concentrated their citations sharply on the same narrow subset.
- The concentration is mostly one shared preference rather than eleven separate ones. A single component explains 68 to 73 percent of the variance across the models’ preference maps, and models from different vendors agree with each other nearly as much as models from the same vendor do.
- So changing engines is not a workaround. Mixing models uniformly removes 28 percent of the excess concentration, and the best weighted mixture the authors could fit still retains 55 percent of it.
- The preference tracks meaning, not wording. Paraphrasing every candidate while preserving its content left the preference map essentially unchanged. That makes surface rewording measured waste in the setting tested. It does not establish which changes of substance move the preference in your favour.
- Scope, plainly: this measures scientific citation on a single topic, it is a preprint, and we have not reproduced it. It is mechanism evidence for the selection layer, not a measurement of your market. Our own Quebec numbers carry the market half.
The version of this question we get most often comes from an operator who has already done the work. The site is fast, the crawlers are allowed through, the pages are indexed, Google sends traffic. Then they ask ChatGPT or Perplexity the question their customers ask, and somebody else gets named. There is now a controlled experiment that isolates the mechanism, and it is not the one most AI-visibility pitches describe.
What the experiment removed
Most published work on AI citations audits the failure everyone already knows about: models inventing references that do not exist. When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models (Alemohammad et al., 2026, a preprint that has not been peer reviewed) asks the question that starts after that one is solved. When every candidate is real, when none is easier to find than the others, and when nothing signals which is prestigious, does the model still play favourites?
What the design removes is the whole point of it. The authors took 120 real papers on one topic. Each prompt showed the model a uniformly random panel of 30 of them, with real titles and abstracts, but with fabricated authors, reassigned years, and hidden venues and citation counts. Retrieval is equalized, since every candidate is already in front of the model. Prestige is gone, since there is no author, journal, or citation count to lean on. Then a hard budget: cite at most 10. Scarcity is what forces a choice, and a choice is what makes a preference visible. Every run was compared against the same model choosing indifferently from the same panel.
Translate that into your market and it reads as an unusually clean control. It is the equivalent of putting your page and nine competitors’ pages in front of a model with the logos off, the domains hidden, and the backlink profiles invisible, then asking which one it quotes.
What it found: concentration that retrieval cannot explain
Every one of the eleven models concentrated. The top decile of papers took between 23.3 and 30.2 percent of all citations, against 15.6 percent under indifferent selection, and papers went persistently uncited across dozens of separate exposures. A formal exchangeability test rejected a favouritism-free selector for every model in the panel.
Eight domain experts judged byte-identical panels under the same citation budget, and their choices were statistically indistinguishable from indifferent selection. Two things are true about that at once, and the honest reading needs both. The experts showed no detectable preference of the kind the models show. And the expert sample was small, 53 prompts with 41 from a single annotator, which is why the authors report a ceiling on any preference they could have missed rather than an estimate of one. What survives is narrow and still substantial: this concentration is a property of the models, not an inevitable feature of choosing among papers under a budget.
One map, not eleven
The models are not each doing their own thing, and that is the finding with the most direct commercial consequence. A single component explains 68 to 73 percent of the variance across their preference maps, and agreement between vendors (the GPT, Gemini, and Claude families) nearly matches agreement within a vendor. Different training runs, different companies, largely the same taste.
That result kills the workaround most people reach for first. If each model had its own idiosyncratic preference, spreading your bets across engines would average the idiosyncrasies out. The authors tested exactly that. Uniform mixing across all eleven models removes 28 percent of the excess concentration. The best cross-fitted weighted mixture they could construct still retains 54.6 percent of it, with a 95 percent confidence interval running from 48.9 to 60.6 percent. Vendor diversity buys you the idiosyncratic part and leaves the shared part standing, and the shared part is most of it.
For an operator, that reframes what “we show up in some AI answers and not others” means. Variation between engines is real, and it is the small half. The large half is a preference they hold in common, which is either an absorbed consensus about what gets cited or a convergent judgment about what a citable source reads like. The paper leaves that question open on purpose, and its results hold under either reading. It does report one partial handle on what the shared map keys on: a methodological-coreness keyword index, built from each paper’s title and abstract, correlates 0.54 with the models’ shared map and only 0.26 with the experts’, with a knowledge-distillation paper at the favoured end and a 3D object-detection paper at the shunned one. The authors file that as a post-hoc correlate rather than the estimate the section rests on, and it accounts for part of the map, not all of it. Read as a direction rather than an instruction, it says the models lean toward what reads as methodologically central.
Which layer your SEO actually reaches
We have described AI search here as layers for a while: eligibility, then selection, then memory. Our GEO field notes lay out the whole model, and the post on citation selection makes the point that no vendor documents the second layer at all. This paper is the first controlled evidence we have seen about what that undocumented layer is doing.
Decided by: robots.txt, your firewall, indexing.
Reached by: technical SEO, completely.
Gets you: into the candidate set. Nothing more.
Decided by: a preference the models largely share.
Reached by: what your content says, not how it is worded.
Gets you: named, or left out of a panel you were already in.
The practical consequence is a scoping rule. Layer-one work is necessary, it is cheap to verify, and it gets you into the candidate set, and we walk that check step by step in the AI SEO mandate. It does not touch layer two. An audit that hands you a green technical scorecard and calls it AI visibility has finished the first layer and left the second one unexamined. Being findable and being preferred are different problems with different work behind them.
The preference reads meaning, not wording
The result that should change a content budget is the paraphrase test. The authors paraphrased all 120 abstracts, hard enough that overlap with the original phrasing nearly vanished (5-gram overlap at or below 0.05) while the content stayed intact. The preference map barely moved: correlation of 0.96 to 1.00 with the original, which is at the ceiling of what the measurement itself can resolve. The two paraphrase conditions that did shift the map were the two that changed what the abstracts actually said, which is a built-in control pointing the same way: content moves the preference and phrasing does not. A crossover test attributed roughly 90 percent of one model’s map variance to the content of the papers.
Read that against what gets sold as AI optimization. Rewriting pages into an “AI-friendly” register, restructuring paragraphs into model-sized chunks, and swapping vocabulary are all surface operations, and this is a controlled measurement of surface operations doing nothing. Google says something compatible from the other direction in its own AI guidance: no special markup, no chunking, no AI writing style.
The inverse is not automatically true. The experiment shows that meaning-preserving edits do not move the preference. It does not show which meaning-changing edits move it in your favour, and nothing here validates any particular tactic. The enrichment hedges we described in the citation-selection post are still hedges. What changes is where the floor sits: if a plan is mostly rewording, there is now a controlled measurement, on scientific abstracts, of that kind of edit moving nothing.
It tightens as AI writing floods the pool
The authors then ran twelve rounds in which model-written papers joined the pool alongside the real ones. The preference itself stayed fixed while the pool grew, and two things happened together. The real papers’ share of all citations fell from everything at the start to 15.6 percent, and each surviving real paper was cited on a rising share of the occasions it appeared: by the last round the originals were cited on 63.4 percent of their appearances, against 39.6 percent for the first generation of model-written work. Part of that trajectory is mechanical, and the authors say so. A larger catalogue under a fixed citation budget concentrates on its own, which lifts the indifferent baseline for top-decile share to 29.1 percent by the last round. The models clear that moving bar at every round, running 31.1 to 40.3 percent. The filter did not wash out with volume. It condensed onto fewer targets.
The scope on that is narrow: this is a fixed preference meeting a growing corpus, not a full feedback ecosystem, because a citation in one round never changes what gets shown in the next. The direction is still uncomfortable for anyone outside the preferred set. As more of the pool is machine-written, the value of already being preferred goes up, and so does the cost of not being.
What this says about your business, and what it does not
The paper measures scientific citation on a single topic, in narrow task formats, choosing among papers already placed in front of the model. It says nothing directly about local businesses, brand mentions, or commercial content. Anyone carrying it into a marketing argument is extrapolating, us included, and that belongs in the open rather than in a footnote.
What transfers is the mechanism: a selection step that sits above retrieval, is largely shared across vendors, keys on content rather than surface, and concentrates. What does not transfer is any number in it. For the market half we use our own measurements instead. The broadest published Quebec study (Observatoire Noos, étude #1) found local SMEs capturing 18 percent or less of model mentions in nine of twelve sector markets, and our own closed-book probe on Quebec roofing had the panel naming a verifiable local business in 7.1 percent of responses. Both are in the post on the memory layer, with their own caveats attached. Two independent lines of evidence pointing the same way is the honest version of this argument. One paper doing all the work is not.
What we do with it
Three things change in how we scope work, and none of them is a promise.
- Measure before arguing. Your position in AI answers is observable: run the questions your customers actually ask across several models and count who gets named. A measurement you can repeat next quarter beats an assurance about how the models work, and it is the honest basis for deciding whether this layer is worth spending on at all.
- Spend on substance, not on surface. If a proposal is mostly rewording, restructuring, and markup, the paraphrase result is a reason to ask what it is expected to change. Budget belongs where the content says something a model would have reason to prefer.
- Do not treat multi-engine coverage as a hedge. Being absent from one engine is weak evidence that another will save you, because the preference is mostly shared. Spread your measurement across engines, not your hopes.
The first of those three is the one you can act on this month without hiring anyone, and it is where we start too: the baseline count of who gets named on your questions, across engines, written down so next quarter has something to compare against. That measurement is the opening move of an SEO audit sprint and the thing we argue from in a consultant engagement, and it is the same posture underneath our GEO work: fix the documented layer completely, treat the inferred layer as inferred, measure rather than promise. What nobody can sell you is a place in the preferred set, because nobody has published the mechanism that would put you there.
The experiment was built to isolate exactly this. The models were not skipping the alternatives because they could not find them. Everything was already in front of them.
