AEOsim Blog
2 Reasons Why the Rich get Richer in LLM Citations
Originally published on LinkedIn by AEOsim.
Table of Contents
Post-hoc rationalization of AI citations is a concern that many in the SEO community have raised [1].
This is closely accompanied by the vague idea of a bias, in some shape or form. However, it can be hard to have a rigorous proof of bias from experiments on closed-source models due to a large number of confounding factors.
Here we detail two different mechanisms to support the intuition behind bias, supported by research, while clarifying how these operate at different steps of the pipeline to compound the Matthew effect [2]. Finally, we talk about what this means for AEO and GEO strategies.
Mechanism 1: The popularity bias is baked into the weights
A NAACL 2025 paper, Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias (Algaba et al., 2025) [3], took 166 ML papers published after GPT-4's knowledge cutoff, stripped out the in-text citations, and asked GPT-4, GPT-4o, and Claude 3.5 to suggest a reference for each one. No search, no RAG. Purely parametric memory: just what the model absorbed during training.
Two key results:
- The models reproduced human citation patterns remarkably well (e.g. title length, publication year, venue, structure of citation network). They had internalized how citations actually work.
- But they did it with a heavy thumb on the scale: a strong pull toward already-highly-cited papers. Among the generated references that turned out to be real (verifiable on Semantic Scholar), the median citation count ran roughly 1,326 higher than the ground-truth references the authors had actually used. The bias held even after controlling for year, title length, venue, and author count.
The idea here is statistical prevalence. Heavily-cited papers show up more often in training data → the model learns them better and recognizes them reliably → it reproduces them more confidently. It is a learned prior which exists before the model does any reasoning. The authors flag that this amplifies the Matthew effect, concentrating attention on a shrinking set of dominant sources.
This is the parametric layer. This layer is inherent to the model and has nothing to do with the retrieval index.
Intuition: The human analogue is buying brand A over brand B simply because you have seen or heard A more times previously, so you remember the name.
Mechanism 2: Reasoning from scratch is expensive, so reputation becomes a shortcut
Evaluating every candidate source on its own merits for every user query is computationally expensive. It costs reasoning tokens, latency, and compute to reason over a large number of sources at inference time, whereas leaning on a known-good source is cheap. So there is a standing economic incentive for a system to substitute reputation for genuine per-query evaluation.
The implication is uncomfortable for smaller players: a brand can win a citation not because its content is the best answer, but because trusting it was the cheapest path to a known-okay answer. On certain queries, better content from an unknown domain just does not get the evaluation budget it would need to win.
To be precise: the NAACL paper does not measure compute cost and does not make this claim. What it establishes is that the same outcome exists even in pure parametric memory. The compute-economics story is a separate mechanism layered on top.
Intuition: Closest human analogue is not wanting to think about difficult and messy topics because this makes you question what you believe, which is hard and uncomfortable.
Why this distinction is important for AEO
These two mechanisms are additive:
- Parametric prior: the model already over-weights popular sources before retrieval runs.
- Retrieval + reranking: the system then has an incentive to lean further on reputation rather than re-evaluate from scratch.
The first is an inherent property of the model, while the second is a design choice.
For AEO/GEO practitioners, this means that the size of the client could be changing your strategy. These mechanisms of bias support the existence of a Pareto frontier between quality and visibility [4]. The shape of this frontier is very likely to be different depending on whether your client is a smaller, newer company or a larger well-known company. You may be able to focus more on quality of content for larger companies while being forced to optimize harder for visibility on smaller brands.
To measure where you actually sit in AI answers, across engines and turns, see Analytika.
Notes
- Post-hoc rationalization is a separate and very likely false claim. Google's RAG attribution patent US20260064780A1 describes a system where retrieved documents shape the answer before any citation decision is made.
- The Matthew effect describes the phenomenon where initial advantages tend to accumulate over time, while initial disadvantages often compound, summarized as "the rich get richer and the poor get poorer."
- Paper: arXiv:2405.15739. Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias.
- A better understanding of why there can be a tradeoff between visibility and quality is facilitated by understanding how the initial Retriever vs Reranker work.
Originally published on LinkedIn by AEOsim. Measure citation and recommendation outcomes with Analytika.