Guide · reviewed August 2026
llms.txt is a plain-text file in Markdown, placed at the root of a domain, that tells a language model what the site is and which pages matter. It is a proposed convention, not a ratified standard.
Its structure is simple: an H1 with the site name, a blockquote summary, then sections of links with a one-line note each. It goes at https://example.com/llms.txt, served as text/plain.
Not the same file
Three files at the root that get confused with each other constantly. They answer different questions, and you want all three.
| File | Question it answers | Read by | Status |
|---|---|---|---|
robots.txt | May you fetch this path? | Every major crawler, including the AI ones | De facto standard, honoured |
sitemap.xml | What pages exist, and when did they change? | Search engines | Standard, honoured |
llms.txt | What is this site, and what should I read first? | Unconfirmed | Proposed convention |
If you only do one of the three, do robots.txt — it is the one with a measurable effect, because a crawler you have not allowed cannot read anything else. Timbaly generates all three per site.
On this page
It is Markdown, deliberately, so that both a model and a person can read it. The convention asks for four things in order:
[Title](url): one-line note.A section named ## Optional has a special meaning: it marks links a model may skip if it is short of context. Everything else is treated as worth reading.
This is the shape of a real file, for an imagined plumbing business. Copy the structure, not the facts.
Two things that make the difference between a useful file and a decorative one. First, the notes carry facts, not adjectives — “4-hour response, 7:00–20:00” is quotable; “fast and reliable service” is not. Second, everything in it must match the pages. A file claiming hours the site contradicts makes both less trustworthy, which is worse than having no file.
An optional companion that contains the full text of your key pages rather than links to them, so a model can read everything in one request. Useful for documentation sites, where the whole corpus is the product. For a small business site it is redundant: the pages are short and already crawlable, and maintaining a second copy of your content is a good way to end up with two versions that disagree.
text/plain. A file that arrives as a styled web page is not the file./contact.robots.txt disallows GPTBot, no llms.txt will change that.Timbaly generates and maintains this file per site, alongside robots.txt with the AI crawlers named explicitly and a sitemap that stays in step with what is published. The rest of the technical layer is here, and the wider GEO reasoning is here.
Questions
A Markdown file at the root of a domain that tells a language model what the site is and which pages matter: an H1 with the name, a blockquote summary, then sections of links with one-line notes. A proposed convention, not a ratified standard.
No. robots.txt controls access and is honoured by every major crawler. llms.txt is about comprehension, for a model already allowed in. One is permission, the other is an index.
Unproven. No major provider has confirmed reading it as a retrieval or ranking input. It is free, reversible, and useful as documentation — publish it, but do not build a strategy on it, and do not pay someone a retainer for it.
https://yourdomain.com/llms.txt, served as text/plain. Not in a subfolder, not behind a redirect.
Yes, per site, and keeps it in step with what is published — along with robots.txt, sitemap.xml and the IndexNow key. You can also edit it directly, including from Claude or ChatGPT over the MCP connector.
Timbaly generates llms.txt, robots.txt and the sitemap for every site it builds, and keeps them current as pages change.