LLM Citations: How ChatGPT, Perplexity, Gemini, Claude and AI Overviews Pick Sources (5 Studies Compared)
How each AI assistant finds and shows the sources behind an answer, what five measurement studies say about which domains get cited, why the percentages disagree, and what an aged domain's inbound links do and do not tell you about citation.
An LLM citation is the link an AI assistant attaches to a sentence in its answer, pointing at the page the sentence came from. Not long ago most assistants had none. Now ChatGPT, Perplexity, Gemini, Claude and Google's AI Overviews all search the live web and cite what they read, and a page that gets cited is a page that gets visited. This guide covers how each one picks and shows its sources, what the published measurement studies say about which domains get cited, why the percentages in those studies disagree with each other, and what any of it means for a site that is new or built on an older domain.
How do LLMs cite sources?
Two different things get called an LLM citation, and it matters which one you mean.
Training is one. A model learns from a large body of text, and Wikipedia, news and reference sites are heavily represented in it. Nothing in that process produces a link. When a model answers from memory it has no source to point at, and any URL it prints is a guess.
Retrieval is the other, and it is where citations come from. The assistant decides the question needs current information, runs one or more web searches, reads the pages that come back, writes an answer, and attaches links to the pages it used. Every assistant below works this way, and every one of them sits on a search index. A page that is not indexed cannot be retrieved, and a page that is not retrieved cannot be cited.
How each assistant picks and shows sources
ChatGPT. OpenAI's help page on searching the web with ChatGPT says "ChatGPT may search the web automatically when your question would benefit from current information", that "Responses that use web search may include citations", and that you can "Select Sources, when available, to view cited sources and other relevant links." The citations appear inline and open the source page when clicked.
Perplexity. Its help centre's how does Perplexity work page says it searches the internet "gathering information from authoritative sources like articles, websites, and journals" and that "Each answer includes numbered citations linking to the original sources". Perplexity was built around citation from the start.
Google AI Overviews and AI Mode. Google's AI features documentation is explicit about eligibility: "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet." It also says "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." Both features can issue several related searches across subtopics to build one response, which Google calls query fan-out, and the doc says the models "identify more supporting web pages, allowing us to display a wider and more diverse set of helpful links associated with the response than with a classic web search."
Gemini. For developers, Google's grounding with Google Search doc says the feature "connects the Gemini model to real-time web content" and returns annotations that "link a text segment" to a source URL. The consumer Gemini app also cites web pages when it draws on Google Search, and Google's eligibility rule above is the one to work from.
Claude. Anthropic's web search tool documentation states "Citations are always enabled for web search", with each citation carrying the source URL, its title, and "Up to 150 characters of the cited content". Developers building on it can pass an allowed_domains or blocked_domains list, which means a Claude-based product can restrict citations to a chosen set of sites before the model ever sees a result.
The common shape is a search index in front of a model. What differs is how many sources each one shows, how much it favours its own properties, and which domains its ranking surfaces first.
Which websites do LLMs cite most?
Five measurement studies, each on its own prompt set and its own definition of share.
| Study | Scope | What it found |
|---|---|---|
| Profound | 680 million citations, August 2024 to June 2025 | Wikipedia is 7.8% of all ChatGPT citations and 47.9% among ChatGPT's top ten sources. Reddit is 6.6% of all Perplexity citations. In AI Overviews, Reddit is 2.2% and YouTube 1.9% of all citations. |
| Semrush | 230,000 prompts, over 100 million citations, 14 July to 12 October 2025, three platforms | "Reddit and Wikipedia are still ChatGPT's two most-cited domains". AI Mode's top five were LinkedIn, YouTube, Reddit, Google and Google Blog. Perplexity's were Reddit, LinkedIn, NIH, Microsoft and Google. |
| Ahrefs | A broad sample of US queries in Google AI Mode | Reddit 17.9%, YouTube 17.8%, google.com 12.5%, Facebook 10.2%, Instagram 5.7%, where "Mention share = a domain's citations as a percentage of the summed citations of the top sources." |
| SE Ranking | 129,000 domains and 216,524 pages across 20 niches, ChatGPT | "Sites with over 32K referring domains are 3.5x more likely to be cited by ChatGPT than those with up to 200 referring domains." Domains with a trust score over 90 "earn almost 4x more citations". |
| Xu, Iqbal and Montgomery, Washington University, arXiv preprint | 55,393 trending queries over 40 days in 2026, Google AI Overviews | AI Overviews appeared on 13.7% of queries and 64.7% of question-form ones. Of the domains cited, "nearly 30% do not appear in those results at all", meaning the first-page results shown beside the overview. |
Read across the rows and the same domain types recur. Reddit and YouTube, where anyone can post. Wikipedia, where anyone can edit but a reference is required. LinkedIn. Government and academic sources such as NIH. News publishers. The engines differ in the mix, with ChatGPT leaning on Wikipedia, Perplexity on Reddit, and Google's features on Google's own properties plus the social platforms, but the set of domain types is stable.
Why the percentages disagree
Wikipedia is 7.8% of ChatGPT citations in one Profound table and 47.9% in another, from the same data. The 7.8% is Wikipedia's share of every citation. The 47.9% is its share of the citations that went to the ten most-cited sources. Ahrefs uses a third denominator, the summed citations of its top sources. Some studies count the share of responses that contain a domain at all, which is a fourth.
None of them is wrong. They are answers to different questions, and a figure quoted without its denominator is not usable. The other thing the top-ten framing hides is how spread out citations are. If the ten biggest sources account for something like a tenth of all citations, then most citations go to everyone else, and that everyone-else is where a normal site lives.
The numbers also move. Semrush tracked Reddit's share of ChatGPT citations at "a steady 3.8% share" from 18 July to 7 August 2025, then "just 0.5%" between 14 and 17 August, "an 86% decline relative to the previous period." A ranking of cited domains is a snapshot of one engine's retrieval settings that month.
Why history and inbound links matter for a new site
Two findings in the table point the same way. SE Ranking found that sites with many referring domains are cited far more often than sites with few. The Washington University measurement study found that nearly 30% of the domains an AI Overview cites are not on the first page of results beside it, which the authors read as "a source selection mechanism distinct from Google's ranking algorithm."
Put together, citation correlates with being referenced by other sites, and it is not the same contest as ranking. A brand-new domain has been referenced by nobody. It can publish, it can earn links, and it will, slowly. What it cannot do is buy the years.
That is the case for looking at an older domain's inbound links before you launch on it. An expired domain that a university, a government body, a newspaper or Wikipedia once linked to carries those links still, and each one was placed by an editor who chose to. Links from a platform where anyone can post in a minute are a different thing, even though those platforms are among the most-cited. We weight links by whether the citing domain chose to publish them, not by how often agents read that domain. What AI citability means for an expired domain explains the reasoning, and AI search citation sources goes deeper on the per-engine data.
To see this on a domain, run it through the agent citability checker, which scores how deeply a domain sits in the part of the web that cited sources link to, and the backlink checker, which lists the referring domains with the notable ones named. In the directory, the Wikipedia and Reddit source hubs group expired domains by which of those sources link to them, so you can compare an editorially placed link set against an open-platform one side by side.
What we can and cannot claim
Agent Citability is a 0 to 100 score. Its methodology is published, with the tier philosophy, example seed domains and the release process on the methodology page. The weights are not published, so a third party cannot reproduce the number, and we say that rather than imply an openness we do not have. Scores are released in versions and can fall between releases as the underlying web graph changes.
What the score is not is a prediction that any engine will cite a given domain. No study in the table tests whether inbound links from cited sources cause citation. They show correlation between link profile and citation, and a selection process that is not plain ranking. We built the metric on the premise that provenance is what those processes reach for, and that premise is stated as a premise.
FAQ
How do LLMs cite sources? By retrieval. The assistant runs a web search, reads the pages returned, writes the answer, and attaches a link to each page it drew from. A model answering from training data alone has no source to cite.
What sources does ChatGPT use? When it searches, the pages its search returns. Across 680 million citations measured by Profound, Wikipedia was the single most-cited domain at 7.8% of the total, with Reddit next. Semrush's later study found the same two domains on top, and also found their share can swing by a large multiple within weeks.
Do AI citations depend on backlinks? They correlate. SE Ranking's study of 129,000 domains found sites with over 32,000 referring domains were 3.5 times more likely to be cited by ChatGPT than sites with up to 200. Correlation is what has been measured. No published study shows the mechanism.
Can I get ChatGPT to cite my website? You can make it possible. The page has to be indexed by the search engine behind the assistant, answer a question directly, and be referenced by sites the engine already trusts. Nothing makes it certain, and the domains that get cited most shift from month to month.
Are AI citations the same as Google rankings? No. Google says a page must be indexed and eligible for a snippet to be cited in AI Overviews, so ranking eligibility is the floor. The Washington University study found nearly 30% of cited domains were not in the first-page results next to the overview.
If the question is which expired domains already carry links from the sources these assistants cite, browse the Revised directory of expired domains with backlinks. Every listing shows the named linking sources, the referring-domain band and the archived history before the name is revealed, and the source hubs let you start from the citing domains you trust most.