How to Find Domains the Sources AI Cites Already Link To

Two steps: work out which sources the assistants lean on for your subject by running your own questions, then find domains those sources already reference. The source families that recur, and why subject fit decides it.

Two steps, and the first one is yours to run rather than to read. Ask the assistant the twenty questions your customers ask, record every citation, and you have the source list for your subject rather than a generic one. Then find domains those sources already reference. The families that recur across every published study are encyclopaedic references, education and government pages, long-lived community threads and trade publications. A domain referenced by one of them inherits nothing automatically: the subject has to match, and that single check disqualifies most high-scoring names. In the Revised directory, as of 14 September 2026, 11,254 of 11,255 listings (99%) name at least one referring source, and 84% pass a low spam screen before anything is listed.

Step 1: work out which sources the assistants cite for your subject

Published cited-domain rankings are a starting point and a bad substitute for your own list. They are measured on somebody else's prompt set, for one engine, in one month.

Do this instead.

  1. Write down twenty questions in the language your customers use. Not keywords. The sentences people type.
  2. Run each one in ChatGPT, Perplexity and Google's AI mode. Use a fresh session each time so prior turns do not steer retrieval.
  3. Record every cited domain, the question, the engine and the date in a spreadsheet.
  4. Repeat monthly. The share numbers move enough that a single run is a snapshot.

Two things fall out. A short list of domains that recur across questions and engines, which is your source list. And a count of how often your own domain appears, which is your baseline. Checking whether AI cites your site sets out the recording method in more detail.

The published studies are the sanity check on your list rather than its replacement. Profound measured 680 million citations and found Wikipedia at 7.8% of all ChatGPT citations and Reddit at 6.6% of Perplexity's. Semrush checked weekly citations for "more than 230,000 prompts over 13 weeks across three LLMs" and found that "Reddit and Linkedin were among the top five most-cited domains on ChatGPT, Google's AI Mode, and Perplexity", with Perplexity's top sources being "Reddit, LinkedIn, NIH, Microsoft, and Google". If your own run looks nothing like that, check your method before you trust it.

Which source families recur?

Across every study the same four types appear, and they behave differently as link sources.

FamilyExamples that recurWhat a link from one is worth
Encyclopaedic referencesWikipediaHigh. Anyone can edit, but a citation is required, so a reference survives review
Education and governmentUniversities, NIH, agency sitesHighest and hardest to obtain. Rarely given and rarely removed
Long-lived community threadsReddit, forums, Q&A sitesLow as a link, high as a mention. Anyone can post one in a minute
Trade publicationsSector press, standards bodiesModerate to high, depending on editorial control over outbound links

The split that matters is whether the citing domain chose to publish the link. An editorial reference is evidence somebody vouched for the subject. A link you could place yourself in sixty seconds is not evidence of anything, even when the platform hosting it is the most-cited domain in AI answers.

Step 2: find domains those sources already reference

You now have a source list. The question becomes which domains those sources link to, and which of those are available.

Three routes, in increasing order of how well they work.

Read the sources' own outbound links. Open the reference pages in your subject and list the domains they cite. Free, slow, and the most accurate reading of intent because you see the sentence the link sits in. Wikipedia article reference lists are the best version of this, and getting backlinks from Wikipedia covers what happens when you try to add one yourself.

Check specific domains you already suspect. Our backlink checker is free and names the notable referring sources rather than only counting them, and the Agent Citability checker will look up any domain we hold a score for.

Start from the source and work backwards. This is what the directory's source hubs are, and it is the step that is slow to do by hand across a whole subject.

The part nobody can do for you is the fit test. A domain referenced by a university for a statistics course is a statistics domain. Publishing anything else on it throws away the reference and, if the mismatch is severe enough, walks into Google's expired domain abuse policy, which covers a domain "purchased and repurposed primarily to manipulate search rankings by hosting content that provides little to no value to users". The same policy's illustrative examples are all subject mismatches: affiliate content on a former government site, casino content on a former school site.

What a reference does and does not transfer

A reference is a fact about the past. Somebody, once, decided this name was worth citing for a subject.

That is the asset, and it is narrower than it sounds. It does not transfer rankings, traffic, an authority score or a citation. It means a page that an assistant's retrieval already reaches for happens to point at a domain you could register. Whether that helps depends on whether the site you build is about the thing the reference was about.

It also decays. Pages get deleted, sites get redesigned, and reference lists get pruned. A domain that has been dead for eight years is quietly losing references the whole time, which is the argument for reading its archived history alongside any score. What AI citability means for an expired domain works through how to read the two together.

Using the Revised directory for this

The directory exists to do step 2 across a whole subject at once. Practically:

  1. Open the source hub, which groups every listing by the sites that reference it.
  2. Pick the source families your own run in step 1 turned up. Wikipedia backlinks held 4,772 listings as of 14 September 2026, .edu backlinks 108 and .gov backlinks 42. Named-source hubs exist for individual universities, agencies and publications too.
  3. In Quality, set a minimum Agent Citability score. It is our published measure of how deeply a domain sits in the part of the web that cited sources link to. Third parties cannot recompute it, it ships as versioned releases, and a score can fall between releases when the pages holding the references change.
  4. Still in Quality, set Spam signal to the low band.
  5. Set Category in Basics to your subject. This is the fit test and it is not optional.
  6. In History, set a minimum Age and restrict Checked to recently verified availability.
  7. Read the Linked by column on the shortlist before revealing anything. The subject those references were about is what you are agreeing to publish in.

Filtering and sorting are free signed in or not. A plan buys reveals, so shortlist first.

Referring domains come from our own web crawl, and every row is screened for spam signals and trademark conflicts before it is listed. What the directory cannot tell you is whether an assistant will cite the domain. It measures provenance, which is an input, not an event.

What does not work

Buying a domain because its Agent Citability score is high and its subject is not yours. The references were about something else.

Placing your own links on the most-cited platforms. The engines cite the thread, not the person who posted in it.

Copying a published top-ten cited-domain list as your source list. One measurement showed an 86% swing in a single domain's share inside a month.

Treating a reference as a transfer. It is evidence about a subject, not a rankings inheritance.

FAQ

How do I find out which sources ChatGPT cites? Run your own questions and record the citations. Twenty questions across three engines, repeated monthly, gives you a source list for your subject that no published ranking can.

Which websites do AI assistants cite most? Wikipedia, Reddit, LinkedIn, YouTube, government and academic sites, and sector press, with the mix differing by engine. Profound measured Wikipedia at 7.8% of all ChatGPT citations across 680 million citations.

Does a link from Wikipedia help me get cited by AI? It is evidence that a reference requirement was met, which is more than most links carry. It is not a guarantee of citation, and adding one yourself is usually reverted.

Can I find expired domains that Wikipedia links to? Yes. That is what a source hub is. The Wikipedia hub held 4,772 listings as of 14 September 2026, each cited from an article in our web crawl.

Does subject matching really matter that much? It is the check that decides most cases. A domain referenced for an unrelated subject gives you a number and no use, and a severe mismatch is the pattern Google's expired domain abuse policy names.

Is Agent Citability an open metric? It is published, not open. The tier philosophy, example seed domains and release process are on the methodology page. The weights are not, so a third party cannot recompute the score, and scores can fall between releases.

If step 2 is the part you would rather not do by hand, browse the Revised directory of expired domains with backlinks. The source hubs group available expired domains by the sites that reference them, with the referring sources named and spam screening on every row, before any name is revealed.

Expired domains, cited by AI

Own a domain ChatGPT and Google already cite

Every listing carries its Agent Citability score, verified referring domains from our web crawl, site history and age, spam filtered out. Reveal a name, register it at any registrar, and put that authority behind your own site.

Availability is re-checked the moment you reveal a name. They go once someone else registers them.