How to use the Robots.txt Generator
- Choose whether search engines can crawl your site.
- Add paths to keep out, like /admin, and your sitemap URL.
- Choose your AI crawler rules, then copy the file to yoursite.com/robots.txt.
What robots.txt does, and doesn't do
robots.txt is a plain text file at the root of your domain that tells crawlers which paths they may fetch. Well-behaved bots follow it. It isn't security: blocked URLs can still appear in search if other sites link to them, and anyone can read the file.
To keep a page out of Google's index, use a noindex robots meta tag instead, and don't block that page in robots.txt, or Google can't see the tag. Generate it with the meta tag generator.
AI crawlers: training vs. search
AI companies run two kinds of bots. Training crawlers, like GPTBot, ClaudeBot, Google-Extended and CCBot, collect pages to train models. Search and assistant crawlers, like OAI-SearchBot, ChatGPT-User, Claude-SearchBot and PerplexityBot, fetch pages to answer a user's question, often with a link back to you.
Many sites block training but allow search, so they still show up in AI answers. This generator lets you choose each group separately. Check that your sitemap and robots.txt are found with the marketing SEO checker.
Frequently asked questions
Where do I put robots.txt?
At the root of the host, so it loads at https://yoursite.com/robots.txt. Each subdomain needs its own file.
Does blocking a page in robots.txt remove it from Google?
No. It stops crawling, not indexing. To remove a page from search, allow crawling and add a noindex meta tag, or remove the page.
Should I block AI crawlers?
It's a business choice. Blocking training crawlers keeps your content out of future models. Blocking search crawlers can keep you out of AI answers and the clicks they send. Many sites block the first group and allow the second.
Do all bots obey robots.txt?
Reputable search engines and the major AI companies say they do. Scrapers and malicious bots often don't, so use your server or CDN to block them.