AI Visibility for a New Website: How to Get Cited by ChatGPT and Google AI Overviews
A new site has no citation history, and AI answer engines pick sources by provenance. What ChatGPT, Perplexity and Google AI Overviews actually cite according to their own docs and the measurement studies, why domains already referenced from Wikipedia, .edu, .gov and Reddit get picked up first, how to measure where you stand, and the limits nobody selling an AI visibility dashboard will tell you.
AI visibility is the share of AI-generated answers in your subject that name your site as a source. It is the same work that gets called AI search optimization or generative engine optimization, and the GEO guide covers the term and where it came from. This guide is narrower. It is about a website that launched recently, has a handful of pages and no history, and wants to know why ChatGPT has never heard of it.
The short answer is that answer engines pick sources by provenance, a new site has none, and the two ways to get some are slow and slower. The longer answer has numbers in it.
What do AI search engines actually cite?
Each engine has published enough to answer this in outline.
Google's documentation on AI features says that "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet", and that "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." Both features may use what Google calls a "query fan-out" technique, "issuing multiple related searches across subtopics and data sources", to develop a response. The candidate pool is the ordinary index. The selection is a retrieval step run several times over it.
OpenAI's crawler documentation says "OAI-SearchBot is used to surface websites in search results in ChatGPT's search features" and that "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links." Training data is a separate crawler, GPTBot. Being in the training set and being cited are different pipelines.
Perplexity's bot documentation uses the same shape: "PerplexityBot is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models."
So all three are search indexes with a language model writing the answer. Which sites the index reaches for is the question, and three measurement studies answer it from different angles. The citation sources article has the platform-by-platform tables. For a new site, three findings matter.
Citations are dispersed. A study of 22,295 AI answers to business-software prompts found 7,058 distinct cited domains across ChatGPT, Perplexity and Google AI Mode, and that "The top 10 domains on each engine cover only 11 to 13% of citations." There is no shortlist of ten sites to get onto. The other 88% is where most citations go.
Selection is not rank. A measurement study of Google AI Overviews from Washington University issued 55,393 queries over 40 days in 2026 and found that cited domains "are more credible than co-displayed first-page results, yet nearly 30% do not appear in those results at all, indicating a source selection mechanism distinct from Google's ranking algorithm." Close to a third of what the engine cited was not on page one for that query, and the engine reached past the ranking for sources it considered more credible.
Reference sources sit at the top. In Profound's analysis of citations from August 2024 to June 2025, Wikipedia alone was 7.8% of all ChatGPT citations and Reddit was 6.6% of all Perplexity citations.
Why does a new website get passed over?
Retrieval has to find you first, and then the model has to prefer you over whatever else came back. The first step is ordinary indexing, and a new site can pass it in days. The second step is where new sites stall, because "credible" in the Washington University finding is not something a page asserts about itself. It is something other sources have said about the domain, by referencing it, over time.
SE Ranking's study of 129,000 domains puts a number on the gap: "Sites with over 32K referring domains are 3.5x more likely to be cited by ChatGPT than those with up to 200 referring domains." A new site has zero referring domains. Every page on it can be quotable and fast, and it still sits outside the layer the engine draws from, because nothing on the web has vouched for the name.
AI visibility for a new site is a provenance problem, and provenance is earned at the rate other people choose to reference you. The content is the easy half. Two ways exist to change the rate. You can earn references on the domain you have, which is slow and is the only way for most sites. Or you can start on a domain that already has them, which is what the rest of this guide is about, with conditions.
Why do domains referenced from Wikipedia, .edu, .gov and Reddit get picked up first?
Every reference a site earns, it earns on a domain. When a site shuts down and the domain expires, the pages that referenced it do not know. A Wikipedia citation or a university library's resource list stays where it is until someone edits it. Of the roughly 12,000 expired domains in the Revised directory, about 5,100 are still referenced from Wikipedia and around 6,100 from Hacker News, years after most of those sites went dark.
| Source | Why engines weight it | What it tells you about the domain | Hub |
|---|---|---|---|
| Wikipedia | Editorially controlled references; the single largest ChatGPT citation source | A Wikipedia editor judged the page a reliable citation for a fact | Expired domains with Wikipedia backlinks |
| .edu | University pages are curated, rarely link out, and are cited by engines for research queries | A department or library listed the site as a resource | Expired domains with .edu backlinks |
| .gov | Government pages link only to what an agency chose to endorse | An official body referenced the site, usually for data or a service | Expired domains with .gov backlinks |
| Community consensus source; largest single Perplexity citation source | People discussed and linked the site without being paid to | Expired domains with Reddit backlinks |
The first three are hard to get and hard to fake, which is why our citability score weights them heavily. Reddit is different. Anyone can post a link there in a minute, so a Reddit reference counts for little in a trust metric, and yet it is the source Perplexity and Google AI Overviews cite most. Both things are true. A Reddit reference is weak evidence of trust and strong evidence that the domain was part of a conversation the engines index. The edu and gov backlinks guide covers where the first two come from.
A domain with those references is already inside the layer the engines draw from. A new site on it starts there instead of outside. That is a head start on the half of AI visibility that cannot be optimised, and it comes with a condition Google has written down. The expired domain abuse policy covers domains repurposed "primarily to manipulate search rankings by hosting content that provides little to no value to users". A domain a university referenced for its climate data is a poor start for a fintech, whatever its score, and a real site on a matched domain is the only version of this that holds up. The aged domain vs new domain comparison covers what to check.
How do you measure AI visibility for a new site?
Three measurements, and they answer different questions.
Start by asking the engines. Write down the twenty questions your site should be the answer to, in the words a customer would use, and run each through ChatGPT search, Perplexity and Google's AI Mode. Record which domains get cited. For a new site the answer will be none of yours. The point is the list of who does get cited, because that list is what the engines consider credible in your subject. Repeat it monthly. It is manual, and it is the only measurement that reflects the answers people actually see.
Then measure your domain's provenance. The citability checker returns a 0 to 100 score for any domain we hold a score for, computed by propagating trust from a tiered seed list of over a thousand sources through the link graph. The backlink checker shows who links to the domain. Run both on your own domain and on the domains the engines cited in step one. The gap between the numbers is the size of the problem.
Last, watch Search Console. Google's documentation says that sites appearing in AI features "are included in the overall search traffic in Search Console" and "reported on in the Performance report, within the 'Web' search type." There is no separate AI Overviews row, so you cannot isolate AI traffic there, but a query that gains impressions while losing clicks is often an AI Overview answering it above your result. The AI Overviews guide goes into reading that pattern.
What can you do on the site itself?
Less than the tools suggest, but not nothing.
Let the search crawlers in. Check robots.txt allows OAI-SearchBot, PerplexityBot and Googlebot. Blocking GPTBot to stay out of training data is a separate choice and does not affect citations. Blocking OAI-SearchBot removes you from ChatGPT search entirely.
Be indexed with a snippet. That is Google's only stated eligibility rule. A noindex or nosnippet directive removes you from the candidate pool before any selection happens.
Answer the question under the heading. Retrieval pulls passages, not pages. A heading in the form of the question and a first paragraph that answers it, with a number and a named source in it, is a passage an engine can quote. The original GEO paper found that adding citations, quotations and statistics improved visibility in generated answers more than any other change it tested, and that keyword density did close to nothing.
Skip the AI-specific files. Google says "You don't need to create new machine readable files, AI text files, or markup to appear in these features." Neither OpenAI's nor Perplexity's documentation mentions llms.txt. If a vendor says otherwise, ask them which engine's documentation they are reading.
All of that makes a page quotable. None of it makes the domain a candidate, which is the order most AI visibility advice gets backwards.
What are the honest limits?
The premise underneath this guide is that engines prefer sources other credible sources already reference. The studies above measure which domains get cited. None of them tests directly whether a reference from a trusted domain predicts that the referenced domain gets cited. We have looked and not found that experiment, and our methodology page says so. The premise is well grounded. It is not proven.
Citations are dispersed and volatile. The 88% figure means there is room for a new source. It also means no single win moves the number much, and engines have changed their sourcing between model releases.
References decay. The pages holding an expired domain's citations get redesigned and deleted. A citability score is a snapshot and can fall between releases. An aged domain is a head start on a clock that is still running, not a permanent asset.
And a domain gets you into the pool for its old subject. If your site is about something else, the references were vouching for a different thing and the engines are unlikely to carry them over. Matching the subject is the whole condition.
FAQ
How do I get cited by ChatGPT with a new website? Allow OAI-SearchBot in robots.txt, get indexed, and write pages that answer a question with a number and a named source in the first paragraph. Beyond that, ChatGPT draws on an index of sites it already has reason to trust, and trust is earned by other sites referencing yours. That takes time, or a domain that already has it.
Does a new domain rank in AI search? It can be indexed within days and cited in principle from then on. In practice the engines skew heavily toward domains with existing references, and a new domain has none.
Is AI visibility the same as GEO or AI search optimization? Yes. They are names for the same work. GEO is the term from the research paper that started the field.
Does an aged domain help with AI visibility? An aged domain with references from sources the engines trust starts inside the layer they draw from. Age on its own does nothing. An old domain nobody cited has no advantage over a new one, and a domain cited for an unrelated subject is a poor fit whatever its score.
How do I measure AI visibility for free? Run your target questions through the engines by hand and record who gets cited. Check your domain's score with a free citability checker and its links with a backlink checker. Watch Search Console for queries gaining impressions and losing clicks.
The domain is the part you cannot rush, so decide it first. The Revised directory of expired domains with backlinks lists domains available to register now, with the referencing sources shown by name on every listing before you reveal it, so you can find one that Wikipedia, a university or a government agency already chose to cite and check whether its subject is yours. Browse it before you write the first page.