Guide · reviewed August 2026

llms.txt, without the hype.

llms.txt is a plain-text file in Markdown, placed at the root of a domain, that tells a language model what the site is and which pages matter. It is a proposed convention, not a ratified standard.

Its structure is simple: an H1 with the site name, a blockquote summary, then sections of links with a one-line note each. It goes at https://example.com/llms.txt, served as text/plain.

The part most guides skip: no major AI provider has confirmed using llms.txt as a retrieval or ranking input. Its effect is unproven. It costs nothing, it takes one deletion to remove, and it doubles as documentation a person can read — so publish it as cheap insurance, and be sceptical of anyone selling it as a service.

Not the same file

Permission, index, and sitemap.

Three files at the root that get confused with each other constantly. They answer different questions, and you want all three.

FileQuestion it answersRead byStatus
robots.txtMay you fetch this path?Every major crawler, including the AI onesDe facto standard, honoured
sitemap.xmlWhat pages exist, and when did they change?Search enginesStandard, honoured
llms.txtWhat is this site, and what should I read first?UnconfirmedProposed convention

If you only do one of the three, do robots.txt — it is the one with a measurable effect, because a crawler you have not allowed cannot read anything else. Timbaly generates all three per site.

The syntax

It is Markdown, deliberately, so that both a model and a person can read it. The convention asks for four things in order:

  1. A single H1 — the name of the site or project. Required.
  2. A blockquote immediately after it — a short summary of what the site is. This is the sentence most likely to be lifted, so write it as carefully as a meta description.
  3. Optional paragraphs of extra context: who it is for, what is out of scope, key facts.
  4. H2 sections containing bullet lists of links, each in the form [Title](url): one-line note.

A section named ## Optional has a special meaning: it marks links a model may skip if it is short of context. Everything else is treated as worth reading.

A complete example

This is the shape of a real file, for an imagined plumbing business. Copy the structure, not the facts.

# Rossi Termoidraulica > Plumbing and heating company in Turin, Italy, founded 1998. > Boiler repair and installation, bathroom renovation, emergency > call-outs within 4 hours across Turin and the first belt. > VAT IT01234567890. # Facts worth quoting: # - Emergency call-out within 4 hours, Monday to Saturday, 7:00-20:00. # - Certified for wall-hung condensing boilers and heat pumps. # - Free written quote before any work begins. # - Areas served: Turin, Moncalieri, Collegno, Rivoli, Nichelino. ## Services - [Boiler repair](https://example.com/boiler-repair): Same-day repair for all major brands, with the call-out fee stated up front. - [Boiler installation](https://example.com/boiler-installation): Condensing boilers and heat pumps, including the subsidy paperwork. - [Emergency plumber](https://example.com/emergency-plumber): 4-hour response, 7:00-20:00, Monday to Saturday. ## About and contact - [About us](https://example.com/about): Family business since 1998, four engineers, all work guaranteed 24 months. - [Contact](https://example.com/contact): Phone, email and the areas covered. ## Optional - [Privacy policy](https://example.com/privacy) - [Cookie policy](https://example.com/cookie-policy)

Two things that make the difference between a useful file and a decorative one. First, the notes carry facts, not adjectives — “4-hour response, 7:00–20:00” is quotable; “fast and reliable service” is not. Second, everything in it must match the pages. A file claiming hours the site contradicts makes both less trustworthy, which is worse than having no file.

What about llms-full.txt?

An optional companion that contains the full text of your key pages rather than links to them, so a model can read everything in one request. Useful for documentation sites, where the whole corpus is the product. For a small business site it is redundant: the pages are short and already crawlable, and maintaining a second copy of your content is a good way to end up with two versions that disagree.

Five common mistakes

  • Serving it as HTML. It must be text/plain. A file that arrives as a styled web page is not the file.
  • Using relative links. Absolute URLs only. A model reading the file out of context cannot resolve /contact.
  • Listing every page. The value is the selection. Forty links with no notes is a sitemap written badly.
  • Letting it go stale. A file describing services you dropped is an active liability. If you cannot keep it current, do not publish it.
  • Expecting it to replace robots.txt. It grants nothing and blocks nothing. If robots.txt disallows GPTBot, no llms.txt will change that.

Timbaly generates and maintains this file per site, alongside robots.txt with the AI crawlers named explicitly and a sitemap that stays in step with what is published. The rest of the technical layer is here, and the wider GEO reasoning is here.

Questions

About the file.

What is llms.txt?

A Markdown file at the root of a domain that tells a language model what the site is and which pages matter: an H1 with the name, a blockquote summary, then sections of links with one-line notes. A proposed convention, not a ratified standard.

Is llms.txt the same as robots.txt?

No. robots.txt controls access and is honoured by every major crawler. llms.txt is about comprehension, for a model already allowed in. One is permission, the other is an index.

Does it actually work?

Unproven. No major provider has confirmed reading it as a retrieval or ranking input. It is free, reversible, and useful as documentation — publish it, but do not build a strategy on it, and do not pay someone a retainer for it.

Where exactly does it go?

https://yourdomain.com/llms.txt, served as text/plain. Not in a subfolder, not behind a redirect.

Does Timbaly write it for me?

Yes, per site, and keeps it in step with what is published — along with robots.txt, sitemap.xml and the IndexNow key. You can also edit it directly, including from Claude or ChatGPT over the MCP connector.

Or let the site write it.

Timbaly generates llms.txt, robots.txt and the sitemap for every site it builds, and keeps them current as pages change.

English Italiano