Do You Need an llms.txt File? What the Evidence Says

No AI company documents reading llms.txt, Google has compared it to the keywords meta tag, and a study of 137,210 domains found 97% of published files were never fetched by anything. The proposal, the evidence against it, the one honest reason to publish one anyway, and where the effort should go instead.

On the evidence available today, no. Publishing an /llms.txt file will not get you cited by ChatGPT, Perplexity, Claude or Google, because no major AI company has documented reading one, and the measurement we have says almost nobody's file is ever fetched by anything.

That is a stronger claim than "it can't hurt", which is the usual advice, so the rest of this is the evidence for it. We publish an llms.txt ourselves, and the reason is not that it works.

What is llms.txt meant to do?

It is a proposal from Jeremy Howard of Answer.AI, published on 3 September 2024 at llmstxt.org. The site describes it as "A proposal to standardise on using an /llms.txt file to provide information to help agents use a website."

The reasoning is sound in the abstract. A web page is built for people, with navigation, adverts and scripts wrapped around the part an assistant wants, and all of that costs tokens. An llms.txt file offers a curated markdown summary of the site with links to the pages that matter, so an agent can orient itself without parsing everything. The format is deliberately minimal: an H1 with the site name is the only required part, and everything after it is optional prose and link sections.

As a documentation convenience for agents that already decided to use your site, that is a reasonable idea. As a way to be discovered or preferred by an answer engine, it is a different claim, and it is the claim being sold.

Does any AI company read it?

None of them says so. This is the part worth checking yourself rather than taking anyone's word for.

OpenAI's crawler documentation describes GPTBot, OAI-SearchBot, ChatGPT-User and OAI-AdsBot, and what each is for. It does not mention llms.txt. Anthropic's crawler page describes ClaudeBot, Claude-User and Claude-SearchBot. It does not mention llms.txt. Perplexity's bot documentation describes PerplexityBot and Perplexity-User. It does not mention llms.txt.

Google is the one vendor that has addressed it directly, twice. Its documentation on AI features says: "You don't need to create new machine readable files, AI text files, or markup to appear in these features. There's also no special schema.org structured data that you need to add."

And in April 2025, Google's John Mueller answered a question about llms.txt in a Reddit thread. Search Engine Journal reported the comment as:

AFAIK none of the AI services have said they're using LLMs.TXT (and you can tell when you look at your server logs that they don't even check for it). To me, it's comparable to the keywords meta tag – this is what a site-owner claims their site is about … (Is the site really like that? well, you can check it. At that point, why not just check the site directly?)

The keywords meta tag comparison is the whole argument in one line. A file in which you describe your own site is an assertion, and an engine that can read the site does not need the assertion.

One thing that confuses people: OpenAI and Anthropic both publish llms.txt files for their own developer documentation. That is them being a good citizen of a documentation convention. It is not evidence that their crawlers consume anyone else's.

Do the files actually get fetched?

Rarely, and that is now measured rather than assumed. Ahrefs published a study of llms.txt files on 15 June 2026, covering 137,210 domains with server-log data from May 2026. Twenty-eight percent of those domains published an llms.txt file. Of those files:

97% of those files received zero traffic in May 2026. Nothing fetched them at all.

Of the requests that did land, 96 percent came from bots, and the largest category was SEO audit tools checking whether the file existed rather than any AI system reading it. AI retrieval bots accounted for around 1.1 percent of requests.

The separate SE Ranking study of ChatGPT citation factors, published 24 November 2025 across 129,000 domains, reached the same place from the other direction: llms.txt files showed negligible or negative association with being cited.

Two independent measurements, one of fetches and one of outcomes, agreeing with what every vendor's documentation implies by omission. That is about as clear as evidence gets in this field.

Why do so many sites have one, then?

Because it costs an hour, because audit tools flag its absence, and because "it can't hurt" is easier to write than "here is the evidence". Adoption is real and growing. Consumption is not.

There is a specific failure mode worth naming. A stale llms.txt is worse than none: it describes a site structure you have since changed, and if anything ever does start reading these files, an inaccurate one is an inaccurate claim about your own site. Anything you publish, you now maintain.

Is there any honest reason to publish one?

Two, and neither is search visibility.

If you run developer documentation, an llms.txt or an .md variant of each page genuinely helps a coding agent read your docs efficiently. That is the use case the proposal was written for, and it works because the agent has already been pointed at you.

And if you want to be ready for a convention that may get adopted, publishing a file you keep accurate is a cheap option to hold. That is our reason. Ours is a short map of the directory, the API and the metrics pages, kept in the route that generates it so it cannot drift from the site. We do not count it as an AI visibility measure and neither should you.

What you should not do is spend a week on it, buy a tool to generate it, or let it displace the work that does move citations.

Using the Revised directory for this

The work that does move citations is provenance: being a domain that sources the engines already cite have themselves referenced. The SE Ranking study found that sites with over 32,000 referring domains were 3.5 times more likely to be cited by ChatGPT than sites with up to 200. No file you write about yourself competes with that.

There are two routes to a reference profile. Earn one, which is slow and is what most sites must do. Or start on a domain that already has one, which is what our directory is for. Concretely:

  • Filter under Quality on a minimum Agent Citability score. Agent Citability is our published measurement of how deeply a domain is referenced by the kinds of sources answer engines cite; it ships as versioned releases and a score can fall between them.
  • Use the source filter or go straight to a hub such as expired domains with Wikipedia backlinks or .edu backlinks, to start from a reference type that cannot be bought anywhere.
  • Read the Linked by column and match the subject. A domain referenced for a subject that is not yours is a poor start whatever its score, and Google's expired domain abuse policy is written about exactly that mismatch.

All filtering and sorting is free, signed in or not. The AI visibility guide for a new website is the long-form version of the argument, and how to get cited by ChatGPT covers the crawler side properly.

FAQ

Does llms.txt help SEO? No. It is not a ranking signal in Google Search and Google's documentation says no new machine-readable files are needed for its AI features either.

Does ChatGPT read llms.txt? OpenAI's crawler documentation does not mention llms.txt, and a May 2026 server-log study found AI retrieval bots accounted for around 1 percent of the requests that reached these files at all.

Should I delete my llms.txt file? No need. Keep it if it is accurate and you will maintain it, delete it if it has gone stale. Just do not count it as AI visibility work.

What is the difference between llms.txt and robots.txt? robots.txt is a long-standing convention that crawlers do honour, and it controls whether they may fetch your pages at all. llms.txt is a 2024 proposal describing your site's contents, with no documented consumer. One is a control, the other is a claim.

What should I do instead? Make sure the search crawlers are allowed in, be indexable, answer the question under the heading with a number and a named source, and work on the references pointing at your domain. That last one is the slow half and the one that shows up in the data.

Expired domains, cited by AI

Own a domain ChatGPT and Google already cite

Every listing carries its Agent Citability score, verified referring domains from our web crawl, site history and age, spam filtered out. Reveal a name, register it at any registrar, and put that authority behind your own site.

Availability is re-checked the moment you reveal a name. They go once someone else registers them.