Should I block AI crawlers in robots.txt?
Last updated 1 September 2026
TLDR
For most small businesses, no — blocking AI crawlers removes you from the pool assistants retrieve from, which costs visibility with buyers who research through them. Blocking makes sense mainly for publishers whose content is the product and who are pursuing licensing. It is a business decision, not a security one.
Key facts
- GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot are the main agents in question.
- Blocking is voluntary compliance; it is not technical enforcement.
- Google-Extended controls AI training use without affecting classic search indexing.
- Many sites block by accident through a plugin or template default.
- You can allow retrieval broadly while disallowing specific private paths.
What am I giving up by blocking?
The chance of being named when someone asks an assistant for options in your category. For a services business, that is the same audience you compete for in search — early, high-intent, and increasingly asking a chatbot first.
You are not gaining protection in exchange. Blocking relies on the crawler honouring the file.
Who should block?
Publishers whose content is the product, businesses with a licensing strategy, and anyone with a specific legal reason. If your website exists to win customers rather than to be the thing sold, blocking works against your goal.
You can also take a middle path: allow the crawlers on public marketing pages, disallow member areas and gated resources.
How do I check my current state?
Read your own robots.txt and look for those user agent names. Then check your server logs to confirm the crawlers behave as the file specifies.
Re-check after any theme or SEO plugin update; defaults change without notice.
Sources
- Published crawler documentation from OpenAI, Anthropic, Perplexity, Google and Common Crawl — User agent names and directive support.
Want this checked on your own site?
The free scan shows your technical health and whether assistants mention you.