---
title: "What Is llms.txt? Format, Example and Does It Work?"
description: "llms.txt is a Markdown file at a site's root that gives AI tools a short summary and a list of key links. The format, a real example and whether it works."
url: https://proxynet.io/blog/what-is-llms-txt
date: 2026-10-05
author: "Acar Diveroli"
category: "AI"
lang: en
---

# What Is llms.txt? Format, Example and Does It Work?

If you run a website, llms.txt has probably reached you in one of two ways. An SEO plugin offers to generate one "so AI can find you", or a Lighthouse report shows "llms.txt does not follow recommendations" under a new Agentic Browsing heading. Then you read that Google Search ignores the file. Both are true, and they stop looking contradictory once you see what the file is for.

This post covers the format, llms-full.txt and Markdown versions of pages, and how the file differs from robots.txt and sitemap.xml. We show part of our own /llms.txt, check what Google, Chrome, OpenAI and Anthropic have said as of 5 October 2026, and end with how to create one.

> **Note: Short answer**
>
> llms.txt is a plain Markdown file at the root of a website (example.com/llms.txt) that gives AI assistants and agents a short summary of the site and a curated list of links to its most useful pages. Proposed by Jeremy Howard in September 2024, it is a convention, not a standard, and it does not control crawling; robots.txt does that. Google says Search ignores llms.txt, while Chrome's Lighthouse checks its format in an experimental category. Its real use is on-demand reading by AI tools, mostly of documentation and product sites.

## What is llms.txt?

llms.txt is a text file written in Markdown that a site places at `/llms.txt`. It tells a large language model (LLM), the kind of AI behind ChatGPT, Claude and Gemini, what the site is and which pages are worth reading. The problem it solves is size: a model can only consider a limited amount of text at once, its context window, and a whole website with its menus and scripts rarely fits.

Jeremy Howard published the proposal on [llmstxt.org](https://llmstxt.org/) on 3 September 2024; a second version, dated 10 August 2026, reflects two years of adoption. The file is meant for inference rather than training: an assistant reads it when a user asks something, not a crawler gathering material for the next model.

It is a convention, with no RFC or W3C specification behind it. Nor is it a rulebook: nothing in the file allows or forbids access to a page.

## What does an llms.txt file look like?

The proposal fixes the order of the parts. Only the first one is required:

1. **An H1 with the name** of the site or project.
2. **A blockquote** with a short summary of the key facts.
3. **Paragraphs or lists without headings,** for details such as how the site is organised.
4. **H2 sections with link lists,** each item a Markdown link with an optional note: `- [Name](url): note`.
5. **A section called `Optional`,** by convention, for links an agent can skip when short of space.

Here is what the file of a small online shop could look like. The shop and its URLs are made up:

```markdown
# Example Garden Supply

> Example Garden Supply sells seeds, garden tools and irrigation parts online and ships to all EU countries.

Every product page shows stock per warehouse. Care guides are written by our own staff.

## Shop

- [Seeds](https://example.com/seeds.md): vegetable and flower seeds, sorted by sowing month
- [Irrigation](https://example.com/irrigation.md): drip kits, timers and spare parts, with compatibility tables

## Help

- [Shipping and returns](https://example.com/shipping.md): delivery times per country, return window and costs
- [Contact](https://example.com/contact): email and phone support hours

## Optional

- [Company history](https://example.com/about.md)
```

The notes matter: an agent picks a link from its name and note alone, so "delivery times per country, return window and costs" beats a bare "Shipping".

## A real example: our own /llms.txt

Our llms.txt at proxynet.io/llms.txt is never edited by hand: a script generates it at every build from the data that renders the pages, so a new blog post cannot be missing. It opens with `# Proxynet` and a four-line blockquote on who we are, what we sell and how billing works. The first section starts like this:

```markdown
## Company

- [Homepage](https://proxynet.io): overview of all proxy types, the network and the panel
```

The last section uses the `Optional` convention:

```markdown
## Optional

- [Proxynet (Deutsch)](https://proxynet.io/de): site home in Deutsch
- [Blog (Deutsch)](https://proxynet.io/de/blog): blog index in Deutsch; every post is also available as plain Markdown by adding .md to its URL
```

English and Turkish pages are listed one by one; the other four languages mirror the English structure, so they sit under Optional. Blog posts are listed with their `.md` address, a plain Markdown copy sent with an `X-Robots-Tag: noindex` header so it never competes with the real page in search. With every post in two languages, the file holds several hundred links; most sites do better with a page or two.

Our first, hand-written version was flagged by PageSpeed Insights with "llms.txt does not follow recommendations". Every link was a bare URL, so the check for a Markdown link found none; it also nested lists under H3 headings and had two placeholder URLs that returned 404. Generating the file fixed all three.

## llms-full.txt and Markdown versions of pages

llms.txt is an index that points to pages. Two companions often appear next to it.

**Markdown versions.** The proposal asks sites to offer a clean Markdown copy of useful pages at the same address with `.md` added (`/guide.html.md`, or `/guide/index.html.md` when the address has no file name). Markdown keeps headings, lists, tables and links but drops navigation and scripts, so the content takes up far less of a context window. Version 2 adds a way to find the copies: a `rel="alternate"` link of type `text/markdown`, in the page's HTML or an HTTP `Link` header. Our blog posts carry that tag.

Behind Cloudflare, [Markdown for Agents](https://developers.cloudflare.com/fundamentals/reference/markdown-for-agents/) does something similar without separate files: when a client sends `Accept: text/markdown`, Cloudflare converts the HTML page to Markdown. It is a switch under **AI Crawl Control** on the Pro, Business and Enterprise plans.

**llms-full.txt.** Not part of the proposal, this is a convention from documentation platforms: the full text of the documentation in one file. Anthropic's developer documentation publishes both, its llms.txt ending with a link to llms-full.txt. It suits an API reference that a coding assistant reads end to end; for a shop, the index is enough.

## llms.txt vs robots.txt vs sitemap.xml

The three files sit side by side at a site's root but answer different questions:

| | robots.txt | sitemap.xml | llms.txt |
|---|---|---|---|
| Question | Which paths may bots crawl? | Which URLs exist? | What should an AI read first? |
| Main reader | Search and AI crawlers | Search engines | AI assistants and agents |
| Format | Plain-text rules | XML | Markdown |
| Status | Standard (RFC 9309) | Sitemaps protocol | Community proposal |
| Controls access? | Asks crawlers to stay out | No | No |
| Scope | Rules only | Every page to index | A short selection |

robots.txt is defined by [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html), which says its rules "are not a form of access authorization": well-behaved crawlers follow them, but nothing enforces them. Our [robots.txt guide](/blog/robots-txt) explains each rule.

A sitemap lists every URL you want search engines to find, up to 50,000 per file under the [Sitemaps protocol](https://www.sitemaps.org/protocol.html); [finding a website's sitemap](/blog/find-website-sitemap) shows where to look. It aims to be complete and carries no notes, while llms.txt is deliberately selective.

## Does llms.txt work?

That depends on who you expect to read it. Here is what each company has said or shipped, checked on 5 October 2026.

**Google Search: no effect.** Google's [AI optimization guide](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide), published on 15 May 2026, says you need no new machine-readable files or Markdown to appear in Search, including its generative AI features. A note added on 15 June 2026 says llms.txt will neither help nor harm visibility or rankings, because "Google Search ignores them".

**Chrome Lighthouse: a format check.** Lighthouse, the auditing tool in Chrome DevTools and PageSpeed Insights, has a new Agentic Browsing category, which [Chrome's scoring page](https://developer.chrome.com/docs/lighthouse/agentic-browsing/scoring) calls experimental (Chrome 150 or later, no 0-100 score). The [llms.txt audit page](https://developer.chrome.com/docs/lighthouse/agentic-browsing/llms-txt) is dated 5 May 2026, and the [source code](https://github.com/GoogleChrome/lighthouse/blob/main/core/audits/agentic/llms-txt.js) shows the whole test: a line starting with `# `, at least one Markdown link and at least 50 characters. A 404 makes it not applicable; a server error fails it. It checks the format, not whether any AI uses the file.

**OpenAI: robots.txt for crawlers.** OpenAI's [crawler documentation](https://developers.openai.com/api/docs/bots) covers OAI-SearchBot (ChatGPT search), GPTBot (training), OAI-AdsBot and ChatGPT-User (actions a user starts) and how to manage them in robots.txt. OpenAI publishes an llms.txt for its own developer docs but has not said its crawlers read other sites' files.

**Anthropic: robots.txt for crawlers.** Anthropic's [help article on crawling](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), updated on 7 April 2026, lists ClaudeBot (training), Claude-User (pages fetched when a user asks) and Claude-SearchBot (search quality), all controlled through robots.txt. It does not mention llms.txt, although Anthropic's developer documentation publishes one.

So none of these companies uses llms.txt as a signal for crawling, ranking or citing, while the AI labs publish it for their own documentation. The use that holds up is on-demand reading: a coding assistant fetches a library's llms.txt to find the right page, or an agent sent to your site reads it instead of guessing from the menu. A text file runs no tracking script, so look for requests to `/llms.txt` in your server logs, not in analytics.

## robots.txt controls AI crawlers, not llms.txt

llms.txt cannot tell AI companies what they may use. Leaving a page out does not hide it, and listing one grants no permission. Crawlers check robots.txt, with one group per user agent:

```text
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /
```

The first two groups ask OpenAI's and Anthropic's training crawlers to stay out. `Google-Extended` controls whether Google may use your content to train Gemini models and for grounding; Google's [crawler documentation](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers) says it does not affect inclusion or ranking in Search. Search crawlers such as OAI-SearchBot and Claude-SearchBot have their own tokens, so a site can stay in AI search while opting out of training. Our own robots.txt explicitly allows GPTBot, ClaudeBot and Google-Extended; either way, the decision belongs in robots.txt.

## How to create an llms.txt file

1. **Choose the pages that answer what a customer would ask an assistant about you:** products or services, pricing, documentation, policies, contact.
2. **Write the H1 and the blockquote,** with the facts a model would otherwise guess: what you offer, to whom and where.
3. **Group the links under H2 headings** as `- [Name](url): note`, with full URLs, and move secondary pages under `## Optional`.
4. **Link to Markdown versions** of the pages if you can offer them.
5. **Upload the file to the root** so `https://yourdomain.com/llms.txt` answers with a 200 status as plain text, then run Lighthouse in Chrome DevTools or PageSpeed Insights and find the llms.txt line under Agentic Browsing.
6. **Keep it current,** or generate it automatically. A stale link sends an agent to a 404.

### Generators, plugins and platforms

- **WordPress.** Some SEO plugins create and update the file. In Yoast SEO it is a switch under **Yoast SEO › Settings › Site features**, in the **AI tools** group ([Yoast's help page](https://yoast.com/help/enable-llmstxt/)), not available on multisite. llmstxt.org also lists AIOSEO.
- **Shopify.** Since a [changelog entry of 28 May 2026](https://shopify.dev/changelog/posts/customize-llmstxt-llms-fulltxt-and-agentsmd), every store serves a default `/agents.md`, and `/llms.txt` and `/llms-full.txt` point to the same content. A `templates/agents.md.liquid` template, or a separate `llms.txt.liquid`, in the theme code editor changes it.
- **Documentation and site builders.** llmstxt.org lists Mintlify, GitBook and Wix among platforms that generate the file, plus plugins for VitePress, Docusaurus and Drupal.
- **Online generators** crawl your site and draft a file. Treat the draft as a start: a crawler cannot know which pages matter to your customers.
- **Your own build.** If the site is built from structured data, as ours is, generate the file in the build so it cannot drift.

## Who reads llms.txt, and how agents should read it

- **API and software documentation.** Coding assistants read docs on demand, often through tools connected over [MCP (Model Context Protocol)](/blog/what-is-mcp). This is where llms.txt started.
- **Shops and booking sites.** Agents that compare products need the facts on your pages; [why AI shopping agents get blocked](/blog/ai-shopping-agents-blocked) covers how sites treat them.
- **Sites people visit through AI browsers,** which read pages much as agents do; see [What Is an AI Browser?](/blog/what-is-an-ai-browser).
- **Teams that build agents.** A site's llms.txt is a cheap first request that says where to go next; [how AI agents work](/blog/how-ai-agents-work) explains the loop.

If you build the agent, read robots.txt first: a link in llms.txt does not override a `Disallow` for your user agent. Treat the file as untrusted input that can carry prompt injection, instructions hidden in content; [safe web access for LLMs](/blog/llm-safe-web-access) covers allowlists and rate limits. Identify your bot with a clear [user agent](/blog/what-is-user-agent). For collecting public pages at scale, our [web crawler](/web-crawler) and [data scraping](/data-scraping) pages show where proxies fit in a crawler that follows these rules.

## Common mistakes

- **Bare URLs instead of Markdown links.** Lighthouse reports "File does not appear to contain any links", as it did for our first version.
- **Using it to block AI training.** Only robots.txt groups carry that request.
- **Pasting the sitemap into it.** Thousands of URLs without notes give an agent no reason to prefer one page.
- **Placeholder or dead links,** such as template variables like `{slug}` or pages that now return 404.
- **A server that answers with HTML.** Some setups return the home page for any unknown path; the response has no Markdown heading, and Lighthouse reports the missing H1.
- **Writing it once and forgetting it.** A stale summary is exactly what an assistant will repeat.

## Decision guide

| Your situation | What to do |
|---|---|
| You want more Google traffic, including AI Overviews | Skip llms.txt for this goal; work on content, crawling and indexing |
| You run API or product documentation | Publish llms.txt with Markdown versions; consider llms-full.txt |
| You run a WordPress site | Switch it on in your SEO plugin, then trim the link list by hand |
| You run a Shopify store | Edit the default agents.md through a theme template |
| You want to keep AI training crawlers out | Add robots.txt groups for GPTBot, ClaudeBot and Google-Extended |
| Lighthouse says "llms.txt does not follow recommendations" | Add an H1, use Markdown links, serve the file as text with a 200 status |

## Frequently asked questions

### Does llms.txt help SEO?

Not in Google Search. Google's AI optimization guide says Search ignores llms.txt files, so they neither help nor harm rankings or AI features. Any benefit comes from AI tools that read the file on demand.

### Do ChatGPT and Claude read llms.txt?

Neither OpenAI nor Anthropic has said that its crawlers use llms.txt; both document robots.txt as the control. An assistant can still open the file when a user or a tool asks it to fetch that address.

### Can llms.txt block AI crawlers or AI training?

No, the file contains no rules. To ask a crawler to stay out, add a group for its user agent token to robots.txt.

### Where do I put llms.txt?

At the root of the domain, so it opens at `https://example.com/llms.txt`. The proposal also allows a subfolder such as `/docs/llms.txt` that covers only the pages below it, useful when you control only part of a domain.

### What is the difference between llms.txt and llms-full.txt?

llms.txt is a short index of links. llms-full.txt holds the full text of the pages in one file; it is a convention from documentation platforms, not part of the proposal.

### Is there an llms.txt validator or checker?

The closest to an official check is the llms.txt audit in Lighthouse. It tests only an H1, at least one Markdown link, a minimum length and a working server response, not whether your links are the right ones.

## Summary

llms.txt is a Markdown index for AI tools: a name, a summary and curated links. It is a proposal, not a standard, and it controls nothing; robots.txt still tells AI crawlers what they may fetch. Google Search ignores it and Lighthouse checks only its format, so its value lies with documentation, shops and services that agents read for a user. Keep yours short and generate it from your site's data. If you build agents or crawlers that read public pages, see our [proxy services](/proxy).
