How to use it
- 1Pick a preset. Cite me, don't train blocks training crawlers and allows AI search and user fetchers; you can change any bot individually.
- 2List the paths every crawler should skip — admin, API, checkout, internal search — and any exceptions under Allow.
- 3Add your sitemap URL so crawlers find new pages faster.
- 4Test a few important URLs with different bots, then copy or download the file and upload it to your site root.
Examples
A SaaS that wants AI citations but not training
The default preset puts GPTBot, ClaudeBot, Google-Extended, CCBot and the other training bots in one Disallow: / group. OAI-SearchBot, Claude-SearchBot and PerplexityBot get their own group that repeats /admin and /api/. Testing GPTBot on /blog/launch-checklist says blocked; OAI-SearchBot on the same path says allowed.
Opening one API path to everyone
Disallow: /api/ blocks your dynamic social images at /api/og/. Add /api/og/ under Allow: it's the longer, more specific match, so it wins for that folder while the rest of /api/ stays blocked. The tester shows which rule decided.
Blocking file types with wildcards
Disallow: /*.pdf$ keeps every PDF out of the crawl. The * matches any characters and the $ anchors the end, so /docs/guide.pdf is blocked but /docs/guide.pdf.html is not.
How the file is built and tested
Blocked AI bots are grouped under one set of User-agent lines with Disallow: /. Explicitly allowed bots get their own group carrying your private paths. Everything else falls under User-agent: * with your Disallow and Allow lists, followed by the sitemap.
The tester follows RFC 9309, the robots.txt standard Google implements: it picks the group with the most specific matching user-agent (falling back to *), then the longest matching path rule; when an Allow and a Disallow match with equal length, Allow wins. OpenAI's and Anthropic's bot names were checked against their crawler documentation on September 30, 2026; the others come from each company's public crawler notes.
Whether AI engines can read you is only half of AI visibility — see how AI SEO tools in our directory track citations.
Limitations
- infoNew AI crawlers appear regularly. A bot that isn't in the list follows your
User-agent: *rules. - inforobots.txt controls crawling, not indexing. A blocked page can still show up in search if other sites link to it — use a
noindextag for pages that must stay out. - infoThe tester checks rules in this file only; it can't see server-side blocks, redirects or login walls.
Sources: RFC 9309 · OpenAI crawlers · Anthropic crawlers · Google crawlers
Questions people ask
Will blocking GPTBot stop my site from appearing in ChatGPT?
add
No — they're separate bots. GPTBot collects training data; OAI-SearchBot builds the index ChatGPT search cites, and ChatGPT-User fetches pages when a user asks. The “Cite me, don't train” preset blocks the training bots and allows the search ones, so you can still be cited.
Does Google-Extended affect my Google rankings?
add
No. Google-Extended only controls whether your content is used for Gemini training and grounding. Googlebot, which crawls for Search, is unaffected — blocking Google-Extended doesn't change rankings or AI Overviews eligibility in Search.
Do AI crawlers actually obey robots.txt?
add
The major, named bots from OpenAI, Anthropic, Google, Apple and Perplexity say they do. robots.txt is a request, not an access control: an unnamed scraper can ignore it. User-triggered fetchers may also behave differently because a person asked for the page. To enforce a block, add a rule at your CDN or firewall.
Why are my private paths repeated under the allowed AI bots?
add
A crawler follows only the most specific group that names it. If OAI-SearchBot has its own group, it ignores everything under User-agent: *, including your Disallow: /admin. Repeating the private paths inside that group keeps them private.
Should I use Crawl-delay?
add
Usually not. Google ignores it entirely, and it only slows the bots that honour it. If a crawler is overloading your server, rate-limit it at the CDN instead. Leave the field empty unless you have a specific bot in mind.
Where does robots.txt go?
add
At the root of each host: https://example.com/robots.txt. A subdomain such as blog.example.com needs its own file. It must be plain text, served with a 200 status, and robots.txt rules are case-sensitive for paths.
