Generative Engine Optimization (GEO): What It Is and Where Domains Fit
GEO is the work of getting cited by answer engines rather than ranked by a results page. What the term means, how ChatGPT, Perplexity and Google AI Overviews actually pick sources according to their own docs, a checklist for a new site, and why a domain's existing citations are the part you cannot rush.
Generative engine optimization (GEO) is the practice of making your content the kind that AI answer engines choose to cite. The term comes from a 2023 paper by researchers at Princeton and IIT Delhi, GEO: Generative Engine Optimization, later published at KDD 2024. They describe generative engines as systems that "satisfy queries by synthesizing information from multiple sources and summarizing them using LLMs", and GEO as "the first novel paradigm to aid content creators in improving their content visibility in generative engine responses".
Three years on there are a few dozen tools selling GEO dashboards. The mechanics underneath have not moved much. This guide covers what GEO is, how the major answer engines say they select sources, what that means for a domain that already has a history, and what to do on a site that does not.
What is generative engine optimization?
Classic SEO optimises for a ranked list. Ten blue links, one query, your position on the page. GEO optimises for a synthesised answer in which a handful of sources are cited by name, and the cited sources get whatever attention the answer sends outward.
The GEO paper measured this with a benchmark of 10,000 queries, GEO-bench, and tested a set of content changes against it. Adding citations, quotations and statistics to a page improved its visibility in generated answers the most, by up to 40% on the paper's metric. Keyword density, the reflex from old SEO, had close to no effect. The authors also found "the efficacy of these strategies varies across domains", which is academic phrasing for "what works for a health answer does not work for a coding answer".
That is the whole practical content of the paper. Write pages that give an engine something to quote and something to attribute. The rest of GEO is about being one of the sources the engine considers in the first place, and that is where the engines' own documentation matters more than any optimisation study.
How do answer engines pick sources?
Each engine has published enough to answer this in outline. None has published enough to answer it in detail, and anyone who tells you the ranking formula for an AI Overview is guessing.
Google AI Overviews and AI Mode. Google's documentation on AI features and your website says that "to be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet", and that "there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary". It describes a "query fan-out" technique that issues "multiple related searches across subtopics and data sources", and states plainly that "you don't need to create new machine readable files, AI text files, or markup to appear in these features". So the candidate pool is the ordinary index, and the selection is a retrieval step over that index.
ChatGPT search. OpenAI documents its crawlers on its bots page. "OAI-SearchBot is used to surface websites in search results in ChatGPT's search features." A separate crawler, GPTBot, "is used to crawl content that may be used in training our generative AI foundation models", and a third agent, ChatGPT-User, fetches pages live when a user's question needs it. Two things follow. ChatGPT search has its own index built from what OAI-SearchBot can reach, so a robots.txt rule that blocks it removes you from the citation pool. And training data and search citations are separate pipelines, so a domain being "in the training set" says nothing about whether it gets cited.
Perplexity. Perplexity's crawler documentation says PerplexityBot is "designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models." Perplexity-User visits pages on demand and may "include a link to the page in its response". Same shape as OpenAI: a search index for citations, a live fetch for the answer, and a stated separation from training.
Every one of them is a retrieval system with a language model behind it. Retrieval picks the candidates. The model writes the answer and decides which candidates get named. Page-level GEO work only touches that second step. Whether you are a candidate at all depends mostly on what the domain earned before the question was typed.
Why are already-cited domains cited again?
Two measurement studies describe the citation record, and they are worth reading together.
A study of 22,295 AI answers to business-software prompts found that the ten most-cited domains accounted for only 11 to 13% of citations on each engine. Around 88% of citations went to everything else. Citations are dispersed. There is no top ten to break into, because the top ten is not where most citations live.
A measurement study of Google AI Overviews from Washington University, which issued 55,393 trending queries over 40 days in 2026, found that of the domains cited in an AI Overview, "nearly 30% do not appear in those results at all", meaning the first-page organic results shown beside the Overview. The authors read this as "a source selection mechanism distinct from Google's ranking algorithm". So close to a third of the sources Google's answer engine reached for were not on page one for that query.
Put those two findings next to the retrieval architecture above and a working model falls out. The engines draw on a wide, dispersed set of sources. Selection is not a straight function of rank. And the sources that get selected tend to be ones the engine already treats as trustworthy, whether through its index, its retrieval features or the model's own priors. A domain that has been referenced by credible sources for years sits inside that trusted layer. A domain registered last month does not, however good its pages are.
I want to be careful here. That is a premise, not a proven mechanism. Nobody outside the engines has run the experiment that would confirm it, and our own citability methodology page says so in as many words. But it is the reason the domain matters to GEO at all, and the reason the checklist below has two halves.
What does GEO mean for a domain's history?
Everything a site earns, it earns on a domain. A link from a university library page or a Wikipedia reference attaches to the name, not to the page it pointed at, and it outlasts that page. When the site shuts down and the domain expires, the references persist until the pages hosting them are edited or removed. The scale shows in our own data. Of the roughly 12,000 expired domains listed in it today, about 5,100 are still referenced from Wikipedia and about 6,000 from Hacker News, years after most of the sites went dark.
For a new site, that means the choice of domain is a GEO decision. A fresh registration starts with zero provenance and earns it at the speed other people choose to cite you. An aged domain that already carries references from sources the engines trust starts inside the layer they draw from. The pages still have to be written and the content still has to be good. But the part of the problem that cannot be rushed is already done.
Two cautions, both from the guide to AI citability on expired domains. First, a domain's provenance is a capacity, not an event. It makes the domain the kind of source engines select from. It does not make any engine cite it. Second, citations decay. The pages that referenced a dead site get redesigned and deleted, and a domain that has been down for years is quietly losing references the whole time. Our own citability score is versioned for that reason and can fall between releases. Treat a score or a source list as a snapshot, and look at the archived history beside it.
The other side of the caution is topical fit. An engine citing a domain for its old subject does not carry over to a new one. A domain that a medical journal referenced for its nutrition data is a poor start for a fintech, whatever the score says, and Google's expired domain abuse policy exists precisely for sites that trade on a domain's old reputation to host unrelated content.
GEO checklist for a new site
The list splits into what the page needs and what the domain needs. Both halves matter. A perfect page on a domain no engine has reason to trust waits a long time for its first citation, and a trusted domain hosting thin pages gets passed over.
Page-level GEO
- Let the search crawlers in. Check robots.txt for OAI-SearchBot, PerplexityBot and Googlebot. Blocking GPTBot to keep content out of training is a separate decision from blocking OAI-SearchBot, and blocking the second one removes you from ChatGPT's citation pool.
- Be indexed with a snippet. Google's eligibility rule for AI Overviews is the index plus snippet eligibility.
noindex,nosnippetandmax-snippet:0all remove you. - Give the engine something to quote. The GEO paper's strongest levers were citations, quotations and statistics. Put the number in the sentence and name where it came from. Keep claims in plain text rather than in images or embedded widgets, which the fetching agents cannot read.
- Answer the question in the first paragraph under the heading. Retrieval systems pull passages, not pages. A heading that is the question and a paragraph that is the answer is a passage an engine can lift.
- Skip the AI-specific files. Google says outright that no new machine-readable files or markup are needed. The same is true for the other engines as far as their documentation goes.
Domain-level GEO
- Measure what the domain already has. The citability checker returns a 0 to 100 score for any domain we hold a score for, and the backlink checker shows who links to it. Run both on the domain you have and on the ones you are considering.
- If you are choosing a domain, weigh provenance alongside the name. An aged domain with references from sources the engines trust is a head start on the half of GEO that cannot be optimised. The directory of expired domains with backlinks shows linking sources by name on every listing, and the source hubs group them by who linked, so you can find domains cited by a university, a newsroom or a reference site before you know which domain it is.
What GEO is not
There is no paid placement in an AI Overview, and the dispersed citation record means there is no small set of gatekeepers to court, so GEO cannot be bought. Nor is it a markup standard, whatever the vendor selling one says. None of the three engines' documentation mentions llms.txt or any AI-specific schema, and the SE Ranking study covered in our citation sources article found no measurable effect from either. Google's own guidance for AI features is a list of ordinary SEO practices, and OpenAI and Perplexity both describe their citation systems as search indexes. Being findable is still the precondition for being cited.
The pieces that go beyond this guide are already written. Which sources AI engines cite most covers the platform-by-platform citation data. How Google AI Overviews work goes deeper on the largest engine. And if the domain half of the checklist is where you are stuck, the Revised directory of expired domains with backlinks is a searchable list of names that credible sources already chose to reference, with the sources shown before you reveal anything.
FAQ
What is the difference between GEO and SEO? SEO optimises for position in a ranked list of links. GEO optimises for being cited in a synthesised answer. The eligibility pool for both is largely the same index, so GEO is closer to an extension of SEO than a replacement.
Is AI search optimization the same as GEO? Yes. AI search optimization, answer engine optimization, AI visibility and LLM optimization are all names for the same work. GEO is the term from the original research paper and the one most tools have adopted.
How do I get cited by ChatGPT? Allow OAI-SearchBot in robots.txt, be indexed, and write pages with quotable facts, named sources and statistics. Beyond that, ChatGPT draws on a search index of sites it already has reason to trust, and a domain's existing references are the slow part of earning that trust.
Does domain age matter for AI citations? Age on its own does not. What matters is what the domain earned while it was live: references from sources the engines treat as trustworthy. An old domain nobody cited has no advantage over a new one.
Can a GEO score go down? Yes. Any measure of citability depends on references that other people control, and those references get deleted, redesigned and re-evaluated. Our citability score is versioned for that reason and a falling score is usually the metric noticing that trust was withdrawn somewhere.