AI Crawler Controls: robots.txt & llms.txt
Decide which AI crawlers may use your site, for model training, AI search or on-demand fetching, generate the robots.txt rules, check an existing file, and publish an llms.txt that points AI tools at your best pages.
Follows RFC 9309: a bot obeys the most specific matching User-agent group (falling back to *), the longest matching rule wins, and Allow wins a tie. Fetching from the browser often fails because sites don’t send CORS headers; paste the file instead.
url | title | note per line, ## Section to start a sectionhttps://yoursite.com/llms.txt. An Optional section marks links that can be skipped when context is short (per the llms.txt proposal).Frequently Asked Questions
How do I block ChatGPT and Claude from training on my site?
Add Disallow rules for their training crawlers, GPTBot (OpenAI) and ClaudeBot (Anthropic), to robots.txt. Blocking training crawlers does not remove you from AI search; those use separate agents such as OAI-SearchBot and Claude-SearchBot.
What is Google-Extended?
A robots.txt token, not a separate crawler. Disallowing it tells Google not to use your content to train or ground Gemini models, while normal Googlebot indexing for Search is unaffected.
Do AI crawlers obey robots.txt?
The major vendors document that their crawlers respect it, but robots.txt is voluntary. Some scrapers ignore it, so use server-side blocking or a WAF if you need enforcement.
What is llms.txt?
A proposed standard: a Markdown file at /llms.txt that gives AI tools a short summary of your site and links to the most useful pages, so they can find clean content without crawling everything.
Does blocking an AI bot affect my Google ranking?
No. Googlebot is separate from Google-Extended and from other AI crawlers. Only block Googlebot itself if you want to leave Google Search.