.htaccess Redirect Generator
Generates .htaccess redirect rules for pages, HTTPS, www, and domain moves.
Build a robots.txt file that tells search engines and AI crawlers what they may crawl. Block private folders, opt out of AI training bots, add your sitemap, and test any URL against the rules.
| User agent | What it is | Rule |
|---|---|---|
Googlebot |
Google Search | |
Bingbot |
Bing and Copilot search | |
DuckDuckBot |
DuckDuckGo | |
Yandex |
Yandex | |
Baiduspider |
Baidu | |
Applebot |
Apple Search and Siri | |
Google-Extended |
Gemini training (control token) | |
Applebot-Extended |
Apple AI training (control token) | |
GPTBot |
OpenAI training | |
OAI-SearchBot |
ChatGPT search | |
ChatGPT-User |
ChatGPT browsing for users | |
ClaudeBot |
Anthropic training | |
PerplexityBot |
Perplexity search | |
CCBot |
Common Crawl | |
Bytespider |
ByteDance | |
meta-externalagent |
Meta AI training |
robots.txt is a plain text file at https://yourdomain/robots.txt. It is made of groups: one or more User-agent lines followed by Allow and Disallow rules.
User-agent: *. It does not combine groups.Allow and a Disallow are equally long, Allow wins. This is the rule set in RFC 9309 and used by Google.* matches any characters, and $ at the end anchors the rule to the end of the URL, so /*.pdf$ blocks PDF files.Disallow: allows everything, and Disallow: / blocks the whole site.Allow all crawlers except in /admin/ and /cart, block GPTBot completely, and list a sitemap:
User-agent: * Disallow: /admin/ Disallow: /cart User-agent: GPTBot Disallow: / Sitemap: https://example.com/sitemap.xml
/blog/post: allowed, no rule matches./admin/users: blocked by Disallow: /admin/./cart-help: blocked, because /cart is a prefix of it. Use /cart$ to block only the exact page./blog/post: blocked by its own group.| User agent | Operator | Purpose |
|---|---|---|
Googlebot | Google Search | |
Bingbot | Microsoft | Bing and Copilot search |
Google-Extended | Controls use of your content for Gemini training (no separate crawler) | |
GPTBot | OpenAI | Training data for OpenAI models |
OAI-SearchBot | OpenAI | ChatGPT search results |
ClaudeBot | Anthropic | Training data for Claude models |
PerplexityBot | Perplexity | Perplexity search index |
CCBot | Common Crawl | Open web dataset used by many AI projects |
Applebot-Extended | Apple | Controls use of your content for Apple AI training |
noindex meta tag, and let the page be crawled so the tag can be seen.Crawl-delay. Bing and Yandex follow it.At the root of the host: https://example.com/robots.txt. Crawlers don't look for it in subfolders, and each subdomain needs its own file.
Add a group for each AI user agent with Disallow: /. In this tool, set GPTBot, ClaudeBot, CCBot, Google-Extended, and the others to Block. Search crawlers such as Googlebot are not affected.
No. It stops crawling, but the URL can still appear in results without a description. To remove a page, allow crawling and add <meta name="robots" content="noindex">.
Yes. Disallow: /Admin doesn't block /admin.
Often used together with the Robots.txt Generator.
Generates .htaccess redirect rules for pages, HTTPS, www, and domain moves.
Creates a web app manifest and the HTML tags to make a site installable.
Percent-encodes and decodes text for URLs.