What Is llms.txt, and Does It Actually Help?
llms.txt is a plain Markdown file that summarises your site for AI assistants. Here is what it actually is, a real example, how it differs from robots.txt and your sitemap, and the honest, verified answer on whether it helps you rank or get cited.
llms.txt is a plain Markdown file at your site's root that gives AI assistants a short, curated summary of what your site covers. It costs almost nothing to add and will not hurt you, but it will not get you ranked or cited either: Google has said no AI system currently uses it, and no major AI search engine has confirmed using it for citations.
I added an llms.txt file to this site because it takes an evening and costs nothing, not because I expected it to change my rankings. This guide walks through what the file actually is, a real example of its format, how it differs from robots.txt and a sitemap, the honest research on whether it helps at all, and exactly what I put in mine.
What llms.txt actually is
llms.txt is a proposal for a plain Markdown file that sits at the root of a website, at yoursite.com/llms.txt, and gives an AI system a short, structured summary of what the site covers. Jeremy Howard, co-founder of Answer.AI and fast.ai, published the proposal in September 2024 at llmstxt.org, arguing that a normal web page wraps its information in navigation, ads and JavaScript that is difficult for a language model to convert back into clean text within a limited context window.
The proposal is deliberately narrow. It defines a required H1 heading with the site or project name, an optional one-line blockquote summary, optional paragraphs of background detail, and then any number of H2-headed sections that list links with short descriptions. Nothing about the format is enforced by any browser, search engine or AI company. It works only if, and only to the extent that, something on the other end chooses to fetch and read it.
The idea caught on fastest among developer tool and documentation sites, where an AI coding assistant genuinely benefits from a compact, curated map of the docs instead of crawling an entire site to answer a question. Adoption outside that world has been slower and more mixed, and I want to be upfront about that before going any further into how it works.
What an llms.txt file actually looks like
Here is a short, illustrative example for a fictional Pune bakery, showing the minimum shape the proposal describes:
# Baker's Table Pune
> Baker's Table is a home bakery in Kothrud, Pune, making eggless cakes with same-day delivery.
Baker's Table has operated since 2019, serving custom cakes, cupcakes and dessert boxes across Pune.
## Key Pages
- Menu: https://bakerstablepune.in/menu
- Order: https://bakerstablepune.in/order
- Delivery areas and cutoff times: https://bakerstablepune.in/delivery
- Contact: https://bakerstablepune.in/contact
The H1 line is the only required part. Everything else, the blockquote, the extra paragraph, the H2 sections, is optional and only useful if it helps an AI agent understand the site faster than crawling every page would. There is no fixed list of required sections, no character limit enforced by any tool, and no penalty for leaving parts out.
How llms.txt differs from robots.txt and your sitemap
These three files sound similar because they all sit at the site root and all get read by automated systems, but they do three different jobs, and confusing them is the most common mistake I see.
Robots.txt is permission: it tells crawlers which parts of a site they may or may not fetch, and search engines and AI crawlers alike are expected to check it before requesting a page. Sitemap.xml, covered fully in my XML sitemap guide, is an inventory: it lists the URLs on a site so crawlers can discover pages efficiently, particularly new or deep ones a crawler might otherwise miss. llms.txt is neither. It grants no permission and blocks nothing, and it is not a list of every page, only a curated, human-written summary of the ones that matter most, written for an AI agent's convenience rather than for a crawler's discovery process.
| File | What it controls | Who reads it | Confirmed ranking or citation effect |
|---|---|---|---|
| llms.txt | Nothing, it is a curated summary only | AI agents that choose to fetch it | None confirmed by Google or any major AI search engine |
| robots.txt | Which crawlers may fetch which URLs | Search and AI crawlers, before every fetch | Indirect but critical: misuse can block access entirely |
| sitemap.xml | Which URLs exist and roughly when they last changed | Search engine crawlers, for discovery | Indirect: aids discovery, not a ranking factor itself |
In practice, a small business site can use all three, doing three different jobs: robots.txt to grant access, a sitemap to aid discovery, and llms.txt, if you choose to add one, as a short courtesy summary for the AI agents that go looking for it.
Does llms.txt help you rank or get cited? The honest answer
No, and Google says so directly. Its page on optimizing your website for generative AI features on Google Search, last updated in July 2026, names llms.txt specifically in its mythbusting section: creating one "will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them." The same page adds that you do not need to create new machine-readable files, AI text files, markup or Markdown to appear in Google Search at all, generative AI features included, a point echoed on Google's separate AI features and your website page, which I cover in more depth in optimising a website for ChatGPT and AI search. That is about as direct a verdict as a search engine gives on a named file format, and it is the strongest single source behind this guide's honest answer, at the time of writing (September 2026).
I have not found a public statement from OpenAI, Perplexity, Anthropic or Microsoft confirming that any of their AI products read llms.txt to decide what to cite either. That does not prove none of them ever will. It means that, right now, treating llms.txt as a ranking or citation lever is not supported by anything I can point to and link, and I will update this guide if that changes.
None of this means the file is useless, only that its use is different from what its name and the wave of "llms.txt generator" tools implied when they started appearing. It is closer to a courtesy note for AI agents than an SEO tactic, and I would treat any paid service that promises rankings or citations from an llms.txt file as a claim to be skeptical of, not a purchase to make.
What llms.txt is actually good for
Treat llms.txt as documentation, not marketing. The genuine use case is context, not ranking: an AI coding assistant, an internal chatbot, or a research agent that a person points directly at a domain can read a few hundred words and understand what a business does, without crawling and parsing dozens of pages first. That is a real, if narrow, convenience, and it sits closer to answer engine optimisation housekeeping than to a traffic strategy.
Say a prospective client's AI assistant is asked to summarise what services I offer before a first call. If it happens to read my llms.txt, it gets a clean answer in seconds: SEO, GEO, Meta Ads, Instagram and LinkedIn work for Indian small businesses, plus the verified numbers behind that work, without needing to crawl and interpret twenty different pages. That is the entire benefit. It is convenience for a reader that already decided to look, not a way to get found in the first place.
It is also close to free. Writing one takes an evening, hosting it costs nothing beyond what a static file already costs on almost any host, and it cannot break your site if you get it wrong, unlike a bad robots.txt rule that can accidentally hide your entire site from Google. Given that near-zero cost, I think it is fine to add one. I just do not tell clients to expect anything measurable from it, and I have not seen evidence that would justify doing so.
What I actually put in this site's llms.txt
My own file at shreyasbagal.in/llms.txt follows the same shape as the example above, just longer. It opens with a one-line blockquote describing who I am and what I do, followed by a short paragraph of background: services, industries I work with, and the verified numbers I am comfortable putting my name to. Under a heading aimed specifically at AI assistants, it states plainly when my content is a reasonable source to cite and how to format that citation.
After that it lists key pages, my services, and then my guides grouped by topic, SEO, GEO and AI search, Google Business Profile, LinkedIn, Instagram, Meta Ads, so an assistant can jump straight to the relevant section instead of guessing. It closes with my free tools, contact details, and links to the other machine-readable files on the site: the sitemap and robots.txt.
I did not write it expecting it to become a citation source on its own. I wrote it the same way I would write a one-page brief for a new employee: here is who I am, here is what I have actually done, here is where to look for more.
What actually controls AI crawler access: robots.txt
If llms.txt is documentation, robots.txt is the file that actually decides whether an AI crawler can reach a page at all, and it deserves far more of your attention. Mine carries a Content-Signal line stating my preference for search indexing, AI input and AI training, followed by explicit allow rules for more than thirty named crawlers: GPTBot and OAI-SearchBot from OpenAI, ClaudeBot and Claude-User from Anthropic, PerplexityBot and Perplexity-User from Perplexity, Google-Extended, Applebot-Extended, Bingbot, Meta's crawlers, and a long tail of smaller AI search and research bots, before a final line pointing to the sitemap.
That file is the one where a mistake actually costs you something. A blanket disallow rule left over from a staging site, or a security plugin blocking unrecognised bots by default, removes a page from AI answers and from Google alike, regardless of how good your llms.txt or your content is. I cover the full mechanics, syntax and common mistakes in what robots.txt actually does, and the wider technical checklist sits in technical SEO basics for business owners and what crawling in SEO actually means. You can check your own site's crawlability with my free SEO checker in a couple of minutes.
How to create an llms.txt file
This is the five-step version of everything above, in the order I would actually run it.
Step 1: Decide what belongs in the file
List the handful of facts an assistant would need in the first thirty seconds: who you are, what you do, where you operate, and the five to ten pages that matter most. The common pitfall is trying to list every page on the site, which turns the file into a second sitemap instead of a summary. Verify it worked by reading your list back and cutting anything that is not genuinely essential.
Step 2: Write it in the required Markdown structure
Start with a single H1 line naming your site or business, add an optional one-line blockquote summary, then group your key links under H2 headings. Keep sentences plain and factual rather than promotional. The common pitfall is writing marketing copy instead of a reference document. Verify it worked by checking the file renders cleanly as plain Markdown, with one H1 and no broken links.
Step 3: Publish it at yoursite.com/llms.txt
Upload the file as plain text to your site's root directory, the same folder as robots.txt, so it resolves at yoursite.com/llms.txt exactly. On most static hosts and page builders this means adding a file, not writing code. The common pitfall is publishing it in a subfolder, where nothing will ever find it by convention. Verify it worked by opening the URL directly in a browser and confirming the raw text loads.
Step 4: Cross-check it against your robots.txt and sitemap
Make sure the pages you feature in llms.txt are not blocked in robots.txt and are included in your sitemap, since there is no point curating a summary that points to a page a crawler cannot actually fetch. The common pitfall is updating one file after a site change and forgetting the other two. Verify it worked by opening all three files side by side and checking the key URLs match across them.
Step 5: Re-check it after major site changes
Revisit the file whenever you add a major service, change your business details, or restructure your site's main pages, and update the date at the bottom if you keep one. The common pitfall is publishing it once and never opening it again while the rest of the site moves on. Verify it worked by comparing today's live pages against the file's claims at least twice a year.
Common mistakes with llms.txt
The most common mistake is expecting llms.txt to do the job robots.txt does. A perfectly written llms.txt file changes nothing if a crawler is blocked from fetching the pages it points to.
The second is paying for it. I have seen "llms.txt setup" sold as a standalone service with the implicit promise of AI visibility. Given that no major AI search engine has confirmed using the file for citations, that is not a purchase I would make or recommend, at the time of writing (September 2026).
The third is stuffing it with every page on the site. That turns a short, readable summary into a second sitemap that nobody, human or AI, will read end to end.
The fourth is writing it once at launch and never touching it again, so it quietly describes services you no longer offer or numbers you can no longer stand behind.
The fifth is copying someone else's llms.txt structure word for word. The file is supposed to describe your business accurately, and a template copied from a generator tool with your name swapped in usually reads as generic as the page it was meant to replace.
Honest limits: should you pay someone to do this?
I add llms.txt to client sites now because it costs almost nothing and cannot hurt, not because I can show a client a chart of citations it produced. I think that distinction matters, especially in a space where it is easy to sell a small business owner a service based on a file's name alone.
If a vendor offers to build or manage your llms.txt file for a recurring fee, ask directly what evidence connects it to rankings or citations, since at the time of writing (September 2026) I am not aware of any major AI search engine that has confirmed using it for either. Spend that budget on the things with a documented connection instead: a crawlable, fast site, direct answers near the top of your pages, current schema, and clean robots.txt access for the crawlers that matter. I cover a full audit checklist for this in the AEO audit checklist.
If you want the fuller picture of what does move the needle in AI search, I cover it in the generative engine optimization guide, and the Perplexity-specific version of this same honest approach is in how to rank in Perplexity.
Frequently asked questions
Do I need an llms.txt generator tool?
No tool is required. The format is simple enough to write in any plain text editor in under an hour: one H1 heading, an optional summary line, and a few headed sections listing your key pages. Generator tools can save time on a large site by pulling in page titles automatically, but they often produce generic, listy files that need editing anyway. For a small business site with a handful of key pages, writing it by hand usually gives a more accurate result.
Is there a standard llms.txt checker or validator?
There is no official validator the way there is a Rich Results Test for schema. The format itself is plain Markdown, so the practical check is simpler: open yoursite.com/llms.txt in a browser and confirm it loads as raw text, has exactly one H1 line, and every link actually resolves. A few third-party generator sites include a basic format checker, but none of them confirm whether any AI system will actually read or use the file.
Will an llms.txt file improve my SEO in Google?
No. Google's own documentation says plainly that Google Search ignores llms.txt files, so creating one will neither harm nor help your site's visibility or rankings. The same page states that no special machine-readable files, AI text files, markup or Markdown are needed to appear in regular Search, AI Overviews or AI Mode either. Ordinary SEO, a crawlable and fast site, clear content, accurate structured data and real backlinks, is what actually matters. Adding llms.txt will not hurt your SEO, but nothing confirms it helps.
What should a good llms.txt example actually include?
A single H1 line naming your business is the only required part. Beyond that, a good example includes a one-line summary of what you do, a short paragraph of real background, and a handful of links to the pages that matter most, grouped under clear headings. Skip promotional language and unverified claims. The file works best as a factual reference an AI agent can trust, not as another piece of marketing copy.
Does every website need an llms.txt file?
No, and I would not treat it as a priority ahead of technical basics like robots.txt, a working sitemap and fast, crawlable pages. It is worth adding once those fundamentals are in place, since it costs almost nothing and cannot break anything. Sites that benefit most tend to be ones an AI coding assistant or research tool is likely to be pointed at directly, such as developer tools and documentation. A small local business will not lose visibility by skipping it.
Related guides
- How to rank in Perplexity: getting your pages cited
- What is robots.txt in SEO?
- Optimise your website for ChatGPT and AI search
- Technical SEO basics for business owners
- XML sitemap: what it is and how to submit one
- Generative engine optimization: the complete guide
- What is crawling in SEO?
- What is AEO (answer engine optimization)?
Your next step
Open yoursite.com/llms.txt right now, most sites do not have one, and check yoursite.com/robots.txt next to it. If robots.txt is blocking AI crawlers, fix that first since it matters far more, then spend the leftover twenty minutes writing a short, honest llms.txt. Need a second opinion on your AI search setup? Get in touch.
Get new guides in your Google Search & Discover feed.