No. 10SEO & growth

robots.txt Generator

Decide, bot by bot, which AI crawlers may read your site — let search and answer engines cite you while keeping training crawlers out — then test any URL against the finished file.

lockRuns in your browser — nothing you type is uploaded.updateLast updated money_offFree, no sign-up
AI crawlers
BotRule
GPTBotOpenAITraining
OAI-SearchBotOpenAIAI search
ChatGPT-UserOpenAIUser fetch
ClaudeBotAnthropicTraining
Claude-SearchBotAnthropicAI search
Claude-UserAnthropicUser fetch
Google-ExtendedGoogleTraining
Applebot-ExtendedAppleTraining
PerplexityBotPerplexityAI search
Perplexity-UserPerplexityUser fetch
Meta-ExternalAgentMetaTraining
CCBotCommon CrawlTraining
BytespiderByteDanceTraining
AmazonbotAmazonAI search
Everyone else (User-agent: *)
robots.txt
# AI crawlers blocked from the whole site
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: CCBot
User-agent: Bytespider
Disallow: /

# AI crawlers explicitly allowed (private paths still excluded)
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Amazonbot
Disallow: /admin
Disallow: /api/
Disallow: /checkout

User-agent: *
Disallow: /admin
Disallow: /api/
Disallow: /checkout

Sitemap: https://example.com/sitemap.xml
Test a URL
blockBlocked for GPTBot — matched group gptbot, rule Disallow: /.
Guide

How to use it

  1. 1
    Pick a preset. Cite me, don't train blocks training crawlers and allows AI search and user fetchers; you can change any bot individually.
  2. 2
    List the paths every crawler should skip — admin, API, checkout, internal search — and any exceptions under Allow.
  3. 3
    Add your sitemap URL so crawlers find new pages faster.
  4. 4
    Test a few important URLs with different bots, then copy or download the file and upload it to your site root.
Worked examples

Examples

A SaaS that wants AI citations but not training

The default preset puts GPTBot, ClaudeBot, Google-Extended, CCBot and the other training bots in one Disallow: / group. OAI-SearchBot, Claude-SearchBot and PerplexityBot get their own group that repeats /admin and /api/. Testing GPTBot on /blog/launch-checklist says blocked; OAI-SearchBot on the same path says allowed.

Opening one API path to everyone

Disallow: /api/ blocks your dynamic social images at /api/og/. Add /api/og/ under Allow: it's the longer, more specific match, so it wins for that folder while the rest of /api/ stays blocked. The tester shows which rule decided.

Blocking file types with wildcards

Disallow: /*.pdf$ keeps every PDF out of the crawl. The * matches any characters and the $ anchors the end, so /docs/guide.pdf is blocked but /docs/guide.pdf.html is not.

Method

How the file is built and tested

Blocked AI bots are grouped under one set of User-agent lines with Disallow: /. Explicitly allowed bots get their own group carrying your private paths. Everything else falls under User-agent: * with your Disallow and Allow lists, followed by the sitemap.

The tester follows RFC 9309, the robots.txt standard Google implements: it picks the group with the most specific matching user-agent (falling back to *), then the longest matching path rule; when an Allow and a Disallow match with equal length, Allow wins. OpenAI's and Anthropic's bot names were checked against their crawler documentation on September 30, 2026; the others come from each company's public crawler notes.

Whether AI engines can read you is only half of AI visibility — see how AI SEO tools in our directory track citations.

Know before you rely on it

Limitations

  • info
    New AI crawlers appear regularly. A bot that isn't in the list follows your User-agent: * rules.
  • info
    robots.txt controls crawling, not indexing. A blocked page can still show up in search if other sites link to it — use a noindex tag for pages that must stay out.
  • info
    The tester checks rules in this file only; it can't see server-side blocks, redirects or login walls.

Sources: RFC 9309 · OpenAI crawlers · Anthropic crawlers · Google crawlers

FAQ

Questions people ask

Will blocking GPTBot stop my site from appearing in ChatGPT?

add

No — they're separate bots. GPTBot collects training data; OAI-SearchBot builds the index ChatGPT search cites, and ChatGPT-User fetches pages when a user asks. The “Cite me, don't train” preset blocks the training bots and allows the search ones, so you can still be cited.

Does Google-Extended affect my Google rankings?

add

No. Google-Extended only controls whether your content is used for Gemini training and grounding. Googlebot, which crawls for Search, is unaffected — blocking Google-Extended doesn't change rankings or AI Overviews eligibility in Search.

Do AI crawlers actually obey robots.txt?

add

The major, named bots from OpenAI, Anthropic, Google, Apple and Perplexity say they do. robots.txt is a request, not an access control: an unnamed scraper can ignore it. User-triggered fetchers may also behave differently because a person asked for the page. To enforce a block, add a rule at your CDN or firewall.

Why are my private paths repeated under the allowed AI bots?

add

A crawler follows only the most specific group that names it. If OAI-SearchBot has its own group, it ignores everything under User-agent: *, including your Disallow: /admin. Repeating the private paths inside that group keeps them private.

Should I use Crawl-delay?

add

Usually not. Google ignores it entirely, and it only slows the bots that honour it. If a crawler is overloading your server, rate-limit it at the CDN instead. Leave the field empty unless you have a specific bot in mind.

Where does robots.txt go?

add

At the root of each host: https://example.com/robots.txt. A subdomain such as blog.example.com needs its own file. It must be plain text, served with a 200 status, and robots.txt rules are case-sensitive for paths.

From the directory

AI SEO Tools worth a look

Launching an AI tool?

List it free on LaunchBoosts — it goes live the minute you submit and enters this week's Launch Race, where the top three win Product of the Week and a dofollow link.