Content Agents

llms.txt: What It Is, Whether It Matters, and How to Generate One

By Ari Ber · September 28, 2026

Category: ai-transformed-workflows

llms.txt: What It Is, Whether It Matters, and How to Generate One

llms.txt is an emerging Markdown file that signals AI crawlers what to index - here's what it actually does, whether you need one, and how to generate it without guessing.

Key takeaways

  1. The problem Teams want to control what AI crawlers index but lack a reliable, enforceable mechanism to do so.

  2. Core insight llms.txt is a useful early signal to AI crawlers, but real control comes from layered measures - clean URL structure, canonical tags, noindex rules, and content in initial HTML rather than JavaScript.

  3. Practical outcome You can generate and deploy an llms.txt file today, then audit your site architecture to give AI crawlers clear, consistent signals across every layer.

We placed an llms.txt file on our domain six months ago. Most AI crawlers ignored it. That is, honestly, the correct place to start this conversation - because if you understand why they ignored it, you understand the entire control problem that llms.txt is trying to solve, and you can make a smarter decision about whether to bother.

robots.txt Was Not Built for This, and AI Crawlers Know It

The naive assumption most teams carry is that robots.txt controls AI crawlers the same way it controls Googlebot. It doesn't. robots.txt is a convention, not a law. It was designed in 1994 for a web where a handful of polite search engine crawlers agreed to check the file before indexing. AI crawlers - GPTBot, ClaudeBot, Applebot, Bytespider - arrived decades later, built by companies with different incentive structures and different definitions of "respect."

The gap this creates is real. A team that sets Disallow: / for GPTBot in robots.txt may feel like they've locked the door. Some crawlers will honor that. Others won't. And because there's no enforcement mechanism, you have no way to know which is which unless you're watching your server logs closely. The result is teams either over-blocking - losing visibility in AI-powered search - or under-blocking, with proprietary content getting indexed anyway.

llms.txt is the emerging attempt to give AI crawlers a cleaner signal. It's not a solution to the enforcement problem. But it's a step toward speaking the right language.

How AI Crawlers Actually Navigate Your Site

Before you can control what AI crawlers see, you need to understand what they check - and in what order. The hierarchy looks like this, roughly: they identify themselves via User-Agent header, check robots.txt for crawl rules, check meta robots tags on individual pages, and sometimes read HTTP headers like X-Robots-Tag for non-HTML content. The word "sometimes" is doing a lot of work in that sentence.

The User-Agent strings to know: GPTBot (OpenAI), anthropic-ai and ClaudeBot (Anthropic), Applebot (Apple), Bytespider (ByteDance), bingbot (Microsoft), Google-Extended (Google's AI training crawler). Each publishes official IP ranges you can verify against. This matters because HUMAN Security research suggests nearly 6% of traffic identifying as an AI crawler is actually spoofed - meaning a bad actor is faking the User-Agent to bypass your rules. If an IP doesn't match the official range for the User-Agent it claims, treat it as spoofed and block at the firewall level.

Meta robots tags give you per-page control. The relevant values: noindex, nofollow, noai, noimageai, nollm, nosnippet. For non-HTML files - PDFs, images, video - the X-Robots-Tag HTTP header carries the same instructions. The gotcha is consistency: not every crawler reads every signal. As documented by Search Engine Land's AI crawler guide, some ignore robots.txt entirely, some skip meta tags. A single rule won't hold. You need layered signals, not one authoritative file.

What llms.txt Actually Is

llms.txt is a Markdown file you place at your root domain - example.com/llms.txt - that describes your site in human-readable format. Think of it as a cover letter to AI crawlers: here's who we are, here's what we do, here's what you're allowed to index. The format is intentionally simple because Markdown is easy for language models to parse.

A minimal llms.txt looks something like this:

# Content Agents

> AI-powered content operations for marketing teams.

Full Name: Content Agents
Short Description: Content workflow tools for modern marketing teams
URL: https://contentagents.dev
Contact Email: [email protected]
Crawl-Agent: GPTBot, ClaudeBot, Bingbot
Crawl-Delay: 10

## Allow
- /blog
- /tools

## Disallow
- /internal
- /drafts

The companion file, llms-full.txt, goes further. If llms.txt is a summary, llms-full.txt is your complete site content rendered as Markdown - every page, every article, structured for direct consumption by a language model. The trade-off is visibility versus exposure: llms-full.txt gives AI systems everything, which is useful if you want to be cited, and risky if your content is your product.

Honest caveat: adoption is still minimal. Most AI crawlers don't check for llms.txt today. It's a hedge against a future where they do, not a working control mechanism right now. Pair it with robots.txt and meta tags for anything that actually needs enforcement.

Blocking or Allowing: The Decision Most Teams Get Backwards

The default instinct is to block. It feels safer. But for most teams with public-facing content, blocking AI crawlers is actively working against your distribution goals. If your content is public and you want visibility in AI-powered search - ChatGPT Search, Perplexity, AI Overviews - you need crawlers to access it. Blocking them removes you from the conversation entirely.

The only cases where blocking makes operational sense: your content is the product. A SaaS company with a knowledge base behind a paywall. A research firm with proprietary reports. A subscription newsletter. In those cases, allowing AI crawlers to index your full content trains competitors' models and undercuts the exclusivity that justifies your price. Block, and do it at multiple layers.

For everyone else - marketing blogs, product documentation, editorial content, free tools - the better position is allow by default, and use llms.txt to signal what you want prioritized. You can't force crawlers to obey, but you can make your preferences clear and hope adoption improves. Teams that block reflexively and then wonder why they're not appearing in AI-generated answers have answered their own question.

On spoofing: verify. Check User-Agent strings against the official IP ranges published by OpenAI, Anthropic, Google, and Microsoft. If the IP doesn't match, block it regardless of what the User-Agent claims. This is a firewall rule, not a robots.txt problem.

The Technical Mistakes We Made Running This in Production

The first mistake was assuming robots.txt would hold uniformly across crawlers. We set rules for GPTBot, assumed the logic applied to the broader category, and then watched server logs show ClaudeBot and Bytespider indexing pages we'd intended to exclude. The fix was adding noai and nollm meta tags to those pages in parallel. One signal is not enough.

The second issue was URL normalization. If your site serves the same content at /page and /page/, or at both www.example.com and example.com, crawlers may index both versions separately. This creates duplicate content signals and dilutes any crawl rules you've set. Set canonical tags, enforce trailing slash consistency in your server config, and make sure your XML sitemap references the canonical version of every URL.

Faceted navigation is a related trap. If your site has filter-driven URLs - /products?color=red&size=large - crawlers will attempt to index every parameter combination. The fix is noindex on filtered pages and canonical tags pointing back to the base URL. Left unmanaged, you'll have hundreds of near-duplicate pages indexed by crawlers that don't know which one to trust.

The fourth mistake costs the most: critical content delivered via JavaScript. Research into AI crawler behavior confirms that Perplexity Sonar Pro, Gemini 2.5 Flash, Claude 4.0 Sonnet, and OpenAI o3 load only HTML. As the behavior pattern goes, content injected by JavaScript after page load is generally invisible to these AI crawlers unless it's also present in the initial HTML. If your main content lives in a JS bundle, crawlers see a shell. Move critical content - headings, body text, structured data, links - into the initial HTML response.

One more constraint worth knowing: as noted in AI crawler research, AI crawlers often operate with more conservative crawl limits than Googlebot, so take note of how deep into your site structure they'll need to go to reach your most important content. Flat site architecture helps. Deeply nested URLs hurt.

How to Generate and Deploy Your llms.txt

Yellow excavator at a construction site, captured from ground level against a clear sky.
Photo by Didgeman on Pixabay

For small sites, the manual approach is fine. Create a Markdown file at /llms.txt, drop in your site name, description, contact email, and crawl rules. The structure above is a working template. Upload it to your root domain and verify by visiting the URL directly in a browser - you should see the raw Markdown or a download prompt.

For sites with hundreds of pages, generate it programmatically. Pull your sitemap, extract URLs by section, and write a script that builds the Allow and Disallow blocks from your content taxonomy. Something like:

function generateLlmsTxt(sitemap, config) {
  const allowedPaths = sitemap
    .filter(url => config.allowList.some(pattern => url.includes(pattern)))
    .map(url => `- ${new URL(url).pathname}`);

  return `# ${config.siteName}

> ${config.description}

URL: ${config.siteUrl}
Contact Email: ${config.contactEmail}

## Allow
${allowedPaths.join('\n')}

## Disallow
${config.disallowList.map(p => `- ${p}`).join('\n')}
`;
}

Wire this to a build step or a cron job that pulls from your CMS, and your llms.txt stays current without manual maintenance.

Add a reference to your llms.txt in your robots.txt file the same way you'd reference a sitemap. And add a <link rel="llms" href="/llms.txt"> tag in your HTML <head> so crawlers that skip the root check can still discover it. Neither of these is formally required today, but both improve discoverability as adoption grows.

We built a free generator at Content Agents that handles this automatically - it reads your sitemap, builds the Markdown, and gives you a file ready to deploy. We're disclosing that here because it's the tool we actually use, and the article exists on the same page as the tool. The generator is free; there's no catch worth hiding.

The Real Control Comes from Site Structure, Not a Single File

Zoom out and the picture is this: llms.txt is a signal, not a lock. It's the right thing to implement now, before crawlers standardize on it, because retroactive cleanup is always harder than early setup. But the actual control you have over what AI systems index comes from clean URL architecture, consistent canonical tags, proper noindex rules, and content in HTML rather than JavaScript.

The tooling landscape for managing AI crawler access is still maturing. Audit your server logs quarterly. Check which crawlers are hitting which URLs. Update your robots.txt and meta tags when your content strategy changes. Keep your XML sitemap current with <lastmod> dates that actually reflect updates - Seer Interactive's analysis of 5,000+ URLs across ChatGPT, Perplexity, and AI Overviews found roughly 65% of AI bot log hits targeted content published or updated within the past year, and around 90% within three years. Recency matters to these systems.

Treat AI crawlers the way you'd treat any search engine: you can't force compliance, but you can give clear signals and make compliance easy. The teams doing that now are building a structural advantage over teams that will have to retrofit it later. llms.txt is ten minutes of work. The site architecture decisions underneath it are worth a quarter of serious attention.

If you're running a WordPress site and want to see how this works in practice, Hostinger Academy's walkthrough offers a clear, hands-on look at implementing llms.txt without requiring a developer background. It's a useful companion to the concepts covered above, especially if you'd rather follow along with a real setup than piece things together from documentation alone.

Frequently Asked Questions

What is llms.txt and what does it do?

llms.txt is a Markdown file you place at your root domain - for example, example.com/llms.txt - that describes your site to AI crawlers. It tells them who you are, what your site covers, and which paths they are allowed or not allowed to index. Think of it as a cover letter for AI systems. It is intentionally simple because Markdown is easy for language models to parse. The companion file llms-full.txt goes further and contains your complete site content rendered as Markdown for direct consumption by a language model.

Do AI crawlers actually respect llms.txt?

Most AI crawlers do not check for llms.txt today - adoption is still minimal. The article is upfront that llms.txt is a hedge against a future where crawlers do support it, not a working enforcement mechanism right now. For anything that actually needs enforcement, you should pair llms.txt with robots.txt rules and meta robots tags on individual pages, since no single signal is honored by every crawler.

Should I block AI crawlers or allow them to index my site?

For most sites with public-facing content - marketing blogs, product documentation, editorial content, free tools - you should allow crawlers by default. Blocking removes you from AI-powered search results like ChatGPT Search and Perplexity entirely. Blocking makes sense only when your content is the product itself, such as a paywalled knowledge base, proprietary research reports, or a subscription newsletter, where AI indexing would undercut the exclusivity that justifies your pricing.

Why is my critical content invisible to AI crawlers?

If your main content is injected by JavaScript after page load, AI crawlers likely cannot see it. Research into AI crawler behavior confirms that crawlers like Perplexity Sonar Pro, Gemini 2.5 Flash, Claude 4.0 Sonnet, and OpenAI o3 load only the initial HTML. Content that lives in a JavaScript bundle is invisible to these systems. The fix is to move critical content - headings, body text, structured data, and links - into the initial HTML response rather than rendering it client-side.

How do I generate an llms.txt file for my site?

For small sites, you can create the Markdown file manually using a simple template that includes your site name, description, contact email, and Allow and Disallow path blocks, then upload it to your root domain. For larger sites with hundreds of pages, the article recommends generating it programmatically by pulling your sitemap, filtering URLs by section, and writing a script that builds the file automatically - then wiring that script to a build step or cron job so the file stays current. You should also reference llms.txt in your robots.txt file and add a link tag in your HTML head to improve discoverability.