llms.txt Explained: How to Control AI Visibility
Last updated on August 13, 2026 at 14:01 PM.llms.txt is a Markdown file in a website's root directory that provides large language models with a curated overview of the most important content – complementing robots.txt, which governs access but says nothing about relevance. Anyone who wants to be cited as a source in ChatGPT, Perplexity or Google AI Overviews needs a technical infrastructure that gives AI systems context: which pages matter, what the brand stands for and where the reliable information lives. Search engines are no longer the only place a brand gets found. Who wants to be cited in ChatGPT, Perplexity or Google AI Overviews needs an infrastructure that speaks to machines – and that work follows rules, not gut feeling. How agentic SEO and GEO make brand visibility measurable across Google and AI answers shows what a data-driven, automated approach to this looks like. This article explains the llms.txt format, presents best practices for LLM optimisation and contextualises how robots.txt, structured data and Generative Engine Optimization work together.

What is llms.txt – and why isn't robots.txt enough?
llms.txt is a community convention, not a formal standard. The specification has existed since 2024, originated with Jeremy Howard and is supported by Anthropic, Perplexity and – observably, though not explicitly confirmed – OpenAI. The crucial difference from robots.txt: robots.txt governs whether a crawler may access content. llms.txt governs what an AI system should consider relevant. Both files coexist and complement each other. Anyone who only maintains robots.txt gives AI systems access but no orientation.
Difference between robots.txt and llms.txt
- robots.txt: Access control at URL level. Defines which paths a crawler may visit and which it may not. Says nothing about relevance, quality or priority.
- llms.txt: Editorial curation for AI systems. Delivers a brand description, prioritises content by function and gives the model context it cannot derive from the crawl alone.
- Interplay: A website needs both. robots.txt opens access; llms.txt prioritises content.
Which platforms support llms.txt
Platform support is unevenly distributed. Anthropic and Perplexity have publicly confirmed the convention. OpenAI shows observable correlations in citations but has not explicitly committed. Google ignores llms.txt and favours its own signal stack of sitemap, robots.txt and structured data. For brands this means: llms.txt is an additional layer on top of schema markup – most effective on the platforms that read it.
| Platform | Confirmation status | Observable effect |
|---|---|---|
| Anthropic (Claude) | Publicly confirmed | Yes – retrieval workflows |
| Perplexity | Publicly confirmed | Yes – page prioritisation |
| OpenAI (ChatGPT) | Not explicitly confirmed | Correlation in citations observable |
| Google (Gemini) | No support | No – favours schema + sitemap |
llms.txt format and structure – what a correct file looks like
A well-structured llms.txt stays under 5 KB, opens with a brand summary as a blockquote and organises links by content function. The file resides at domain.tld/llms.txt and is formatted in Markdown. Anyone who misuses the file as a sitemap replacement – 200 URLs without context – defeats the purpose. AI systems treat an overloaded llms.txt as noise and ignore it.
Required components and Markdown formatting
- Blockquote with brand description: A paragraph summarising the brand, its offering and its positioning in two to three sentences. This summary is demonstrably used by Claude and Perplexity for retrieval decisions.
- Sections by content function: Core Pages, Documentation, Research, Policies – each with Markdown links and a one-line description per URL.
- Contact information: Licensing, press, API access – so AI systems can point to the right contact when uncertain.
Common mistakes when creating the file
Three common mistakes are documented and avoidable: misusing the file as a sitemap, using generic template text without brand customisation, and creating contradictions with robots.txt when both files are maintained by different teams.
| Mistake | Impact | Solution |
|---|---|---|
| Too many URLs (more than 50) | AI ignores the file | Limit to 15–30 curated links |
| Generic text | Low trust signal | Write an individual brand description |
| Contradiction with robots.txt | Inconsistent signals | Maintain both files in coordination |
Identifying LLM crawlers – GPTBot, ClaudeBot and PerplexityBot at a glance
Between May 2024 and May 2025, total AI and search crawler traffic rose by 18 per cent according to a Cloudflare analysis. GPTBot grew by 305 per cent. PerplexityBot recorded growth of over 157,000 per cent – albeit from a very low baseline, which explains the extreme percentage. These figures come from a fixed customer cohort set and demonstrate: AI crawlers have become a substantial share of bot traffic. Anyone who does not know which bots visit their website cannot make an informed decision about access and visibility.
The most important AI crawlers and their user agents
- GPTBot (OpenAI): Share of crawler traffic rose from 2.2 per cent to 7.7 per cent. Used for LLM training and ChatGPT retrieval.
- ClaudeBot (Anthropic): Share fell from 11.7 per cent to 5.4 per cent – possibly a relative decline driven by the strong growth of other bots.
- Meta-ExternalAgent: New entrant with a 19 per cent share of pure AI crawler traffic. Purpose: data collection for Meta LLMs.
- PerplexityBot: Strongest relative growth of all measured bots (from a very low baseline). Delivers real-time web data for Perplexity's AI answers.
robots.txt configuration for AI crawlers
Among the top 10,000 domains, GPTBot is simultaneously the most frequently blocked and the most frequently explicitly allowed bot: 312 domains block it, 61 domains explicitly allow it. The strategic decision is binary: blocking protects content from unwanted training. Allowing increases the probability of being cited in AI answers. Anyone who wants to combine protection and visibility needs a differentiated robots.txt with partial allow rules for specific paths.
| Crawler | Growth May 2024–2025 | Primary purpose |
|---|---|---|
| GPTBot | +305 % | LLM training and ChatGPT retrieval |
| PerplexityBot | +157,000 % (from low baseline) | Real-time web data for AI answers |
| Googlebot | +96 % | Indexing + AI Overviews |
Optimising your website for ChatGPT and AI search – the technical foundation
LLM optimisation (LLMO) encompasses ensuring crawlability, providing structured data and preparing content so that AI systems recognise it as a citable source. None of these layers works in isolation. A perfect llms.txt is useless if robots.txt blocks access. Schema markup is useless if the content itself delivers no clear answers.
Structured data for LLMs – Schema.org as a semantic layer
Schema.org markup delivers the entity clarity that LLMs use for citation decisions. Organization, Article, FAQPage and HowTo form the minimum markup for AI visibility. JSON-LD is the preferred format because it is decoupled from HTML rendering and can be parsed directly by crawlers. Schema does not replace good content – but without schema the model lacks confirmation that a page actually covers what it claims to cover.
PDF normalisation and content formats
PDFs are difficult for LLMs to parse. Tables, columns and embedded graphics are lost during text extraction. The solution: provide an HTML version for every PDF and link to the HTML version in the llms.txt. Within content, a clear heading hierarchy, short paragraphs and explicit definitions at the first occurrence of a technical term measurably improve extraction by RAG systems.
The way agent-supported content operations handle repurposing, executive ghostwriting and quality assurance in a consistent brand voice lays out how that balance is kept.
A documented LLM optimisation strategy makes technical requirements and content priorities plannable. Anyone who does not want to build this capability in-house can develop it with a specialised content marketing agency such as Crispy Content®.
Optimising for AI Overviews – why AI visibility is becoming mandatory
AI Overviews reduce organic clicks significantly. According to an Ahrefs analysis of 300,000 keywords, the click-through rate for top-ranking results fell markedly between December 2023 and December 2025 – of 100 historical clicks, only around 42 remain when an AI Overview is active. Anyone who does not appear as a source in AI answers loses traffic regardless of their classic ranking.
GEO – Generative Engine Optimization as a new discipline
Generative Engine Optimization (GEO) is the optimisation for AI-generated answers on platforms such as ChatGPT, Perplexity and Google AI Overviews. A study by Aggarwal et al. shows: GEO can increase AI visibility by up to 40 per cent. The core factors are entity clarity, citation signals from earned media, consistent brand mentions and source credibility. AI systems demonstrably favour third-party sources over brand-owned content – making digital PR a direct GEO lever.
Difference between classic SEO and LLM SEO
Classic SEO optimises for a position in the SERPs. LLM SEO optimises for a citation in AI answers. Both require technical excellence, but LLM SEO prioritises answerability over keyword density. A paragraph that answers a question directly, factually and in a self-contained way has a higher chance of being cited than a page with perfect keyword distribution but vague statements.
| Criterion | Classic SEO | LLM SEO / GEO |
|---|---|---|
| Goal | Position 1–10 in SERPs | Citation in AI answers |
| Success metric | Rankings, CTR | Citation rate, brand mentions |
| Technical foundation | Sitemap, meta tags | llms.txt, schema, Markdown structure |
RAG optimisation – preparing content for Retrieval-Augmented Generation
RAG systems (Retrieval-Augmented Generation) search external sources in real time to substantiate LLM answers. To be selected as a source, content must be self-contained and factually robust. This distinguishes RAG-optimised content from classic web content that relies on navigation and context within a website. A RAG system extracts a paragraph from its context and must still be able to use it correctly.
- Snippet-ready paragraphs: Each section answers a question completely without requiring the reader to know the rest of the page.
- Explicit conditions: Spell out assumptions – "for B2B companies with more than 50 employees" rather than "in this context".
- Numbers and scenarios: Concrete data rather than vague assessments. A RAG system is more likely to cite "42 per cent fewer clicks" than "significantly fewer clicks".
- Clear pronoun references: No ambiguity with "it", "this" or "that". When in doubt, repeat the referent.
Measuring and managing AI visibility – metrics for LLM optimisation
AI visibility can be operationalised via three metrics: citation rate (how often the brand is cited in AI answers), brand mentions (how often the brand is named, even without a link) and referral traffic from AI platforms (measurable via GA4 attribution). Brands with a well-curated llms.txt see a measurable uplift in citations – especially on Anthropic and Perplexity, where the correlation between llms.txt quality and citation frequency is strongest.
The monitoring stack for LLM optimisation consists of an llms.txt audit (structure, coverage, consistency with robots.txt), schema validation (JSON-LD error-free status, entity coverage) and citation tracking (cross-platform measurement of citation frequency). Industry analysts such as Gartner forecast that AI-driven search visitors could account for a significant share of previous organic traffic by 2028. Anyone who builds this monitoring stack today will have a data foundation in 12 months that supports strategic decisions.
A message only stands out when the content behind it is exceptional – everything else drowns in the noise of the competition. For anyone who would rather focus on their brand, their message and their business than on the technical groundwork, the overview of AI-driven visibility and content services from Crispy Content® covers where that division of labour begins.
The future of LLM optimisation – trends through 2027
Three developments will shape the next 12 months: formal standardisation of llms.txt, tooling maturation and potential Google support. None of these developments is certain, but all are probable enough to base decisions on today.
IETF discussions on standardisation are under way. Regulatory pressure on AI training is growing in the EU and the US. A formal specification would give compliance teams in regulated industries – financial services, healthcare, legal – the basis to approve llms.txt. Until then, the convention remains stable enough for production use but too informal for some enterprise approval processes.
Tooling is reaching sitemap maturity: dedicated llms.txt generators, validators and publishing tools will reach the maturity level of today's sitemap tools by 2026/2027. The biggest open question remains Google's position. If Google formally supports llms.txt, the convention becomes universal. If Google decides against it, it remains a signal for Anthropic, Perplexity and OpenAI – which is already sufficient for most brands.
The enterprise trend with the strongest growth: brand governance via llms.txt. Large brands use the file to steer AI systems towards canonical product descriptions and reduce the rate at which models fabricate or distort product claims.
llms.txt, GEO and LLM SEO – the infrastructure for AI visibility is ready now
The technical infrastructure for AI visibility consists of robots.txt for access control, llms.txt for editorial curation and structured data for semantic clarity. These layers complement each other. Brands that implement all layers now secure a measurable advantage in citations across ChatGPT, Perplexity and Google AI Overviews – before the convention becomes an industry standard and the early-adopter lead shrinks. The mechanics follow a familiar pattern: whoever had a sitemap first was indexed first. Whoever has an llms.txt first will be cited first.
Sources
- Cloudflare (2025): From Googlebot to GPTBot: Who's Crawling Your Site in 2025. URL: https://blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025/ (accessed 13 August 2026).
- Semrush (2025): We Studied the Impact of AI Search on SEO Traffic. URL: https://www.semrush.com/blog/ai-search-seo-traffic-study/ (accessed 13 August 2026).
- Ahrefs (2026): Update: AI Overviews Reduce Clicks. URL: https://ahrefs.com/blog/ai-overviews-reduce-clicks-update/ (accessed 13 August 2026).
- Search Engine Land (2026): Mastering Generative Engine Optimization in 2026: Full Guide. URL: https://searchengineland.com/mastering-generative-engine-optimization-in-2026-full-guide-469142 (accessed 13 August 2026).
- Omnibound AI (2026): Generative Engine Optimization Statistics (2026): 60+ Data Points. URL: https://www.omnibound.ai/blog/generative-engine-optimization-statistics (accessed 13 August 2026).
- Grupa Insight (2026): Structured Data in 2026: Schema.org, AI Search and E-E-A-T. URL: https://grupainsight.com/articles/structured-data-in-the-era-of-ai-search-how-schema-org-strengthens-e-e-a-t-and-dominates-ai-overviews-sge-and-zero-click-results (accessed 13 August 2026).
- Search Engine Land (2025): Meet llms.txt, a Proposed Standard for AI Website Content. URL: https://searchengineland.com/llms-txt-proposed-standard-453676 (accessed 13 August 2026).
- OpenAI (2025): Overview of OpenAI Crawlers. URL: https://developers.openai.com/api/docs/bots (accessed 13 August 2026).
Gerrit Grunert
Gerrit Grunert is the founder and CEO of Crispy Content®. In 2019, he published his book "Methodical Content Marketing" published by Springer Gabler, as well as the series of online courses "Making Content." In his free time, Gerrit is a passionate guitar collector, likes reading books by Stefan Zweig, and listening to music from the day before yesterday.