Robots.txt Generator

Build a robots.txt file that tells search engines and AI crawlers what they may crawl. Block private folders, opt out of AI training bots, add your sitemap, and test any URL against the rules.

Updated
Built in your browser. Nothing is sent to your site.
Default for all crawlers
Specific crawlers
User agentWhat it isRule
Googlebot Google Search
Bingbot Bing and Copilot search
DuckDuckBot DuckDuckGo
Yandex Yandex
Baiduspider Baidu
Applebot Apple Search and Siri
Google-Extended Gemini training (control token)
Applebot-Extended Apple AI training (control token)
GPTBot OpenAI training
OAI-SearchBot ChatGPT search
ChatGPT-User ChatGPT browsing for users
ClaudeBot Anthropic training
PerplexityBot Perplexity search
CCBot Common Crawl
Bytespider ByteDance
meta-externalagent Meta AI training
robots.txt

Test a URL

How to use the Robots.txt Generator

  1. Choose the Default for all crawlers, then list Disallowed paths and Allowed paths, one per line, and your Sitemap URL.
  2. Under Specific crawlers, set any bot to follow the default, allow everything, or block everything, for example to opt out of AI training crawlers.
  3. Copy or download robots.txt and upload it to the root of your site. Use Test a URL to check whether a crawler may fetch a given path.

How it works

robots.txt is a plain text file at https://yourdomain/robots.txt. It is made of groups: one or more User-agent lines followed by Allow and Disallow rules.

  • A crawler uses the most specific group that names it, and falls back to User-agent: *. It does not combine groups.
  • Within a group, the rule with the longest matching path wins. When an Allow and a Disallow are equally long, Allow wins. This is the rule set in RFC 9309 and used by Google.
  • * matches any characters, and $ at the end anchors the rule to the end of the URL, so /*.pdf$ blocks PDF files.
  • An empty Disallow: allows everything, and Disallow: / blocks the whole site.

Examples

Allow all crawlers except in /admin/ and /cart, block GPTBot completely, and list a sitemap:

User-agent: *
Disallow: /admin/
Disallow: /cart

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap.xml
  • Googlebot on /blog/post: allowed, no rule matches.
  • Googlebot on /admin/users: blocked by Disallow: /admin/.
  • Googlebot on /cart-help: blocked, because /cart is a prefix of it. Use /cart$ to block only the exact page.
  • GPTBot on /blog/post: blocked by its own group.

Common crawler user agents

User agentOperatorPurpose
GooglebotGoogleGoogle Search
BingbotMicrosoftBing and Copilot search
Google-ExtendedGoogleControls use of your content for Gemini training (no separate crawler)
GPTBotOpenAITraining data for OpenAI models
OAI-SearchBotOpenAIChatGPT search results
ClaudeBotAnthropicTraining data for Claude models
PerplexityBotPerplexityPerplexity search index
CCBotCommon CrawlOpen web dataset used by many AI projects
Applebot-ExtendedAppleControls use of your content for Apple AI training

Limitations

  • robots.txt is a request, not a lock. Well-behaved crawlers follow it; scrapers can ignore it. Protect private content with a password.
  • Blocking a page from crawling doesn't remove it from search results if other sites link to it. Use a noindex meta tag, and let the page be crawled so the tag can be seen.
  • Google ignores Crawl-delay. Bing and Yandex follow it.
  • Crawler names change over time. Check each operator's documentation for the current user agent.

Frequently asked questions

Where do I put robots.txt?

At the root of the host: https://example.com/robots.txt. Crawlers don't look for it in subfolders, and each subdomain needs its own file.

How do I block AI crawlers?

Add a group for each AI user agent with Disallow: /. In this tool, set GPTBot, ClaudeBot, CCBot, Google-Extended, and the others to Block. Search crawlers such as Googlebot are not affected.

Does Disallow remove a page from Google?

No. It stops crawling, but the URL can still appear in results without a description. To remove a page, allow crawling and add <meta name="robots" content="noindex">.

Are paths case-sensitive?

Yes. Disallow: /Admin doesn't block /admin.

Often used together with the Robots.txt Generator.