Skip to main content

AI Robots.txt Checker

Check if your website allows AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) to access your content. Free, instant results.

Why AI Crawler Access Matters

AI search engines like ChatGPT, Perplexity, Claude, and Google Gemini use web crawlers to index content. If your robots.txt blocks these crawlers, your content won't appear in AI-generated answers, meaning you're invisible to a growing segment of search users.

Which AI Crawlers Should You Allow?

If you only allow a few, allow the crawlers that decide whether an engine can cite you: OAI-SearchBot (ChatGPT), Claude-SearchBot, PerplexityBot, Bingbot (Copilot), and Google-Extended (Gemini). Google-Extended is the surprising one Google documents it as controlling both training and grounding, and grounding is how Gemini cites you, so blocking it does cost citation. GPTBot is documented as the training crawler and ClaudeBot as feeding model training, so blocking those is a separate decision about model training and does not by itself remove you from ChatGPT or Claude answers.

Frequently asked questions

Can AI bots read my site?
Enter your domain and this free checker parses your robots.txt and reports which AI crawlers (GPTBot, ClaudeBot, PerplexityBot, and others) are allowed or blocked from reaching your content.
Which AI crawlers should I allow?
For citations specifically, the ones that matter are OAI-SearchBot (ChatGPT), Claude-SearchBot (Claude), PerplexityBot (Perplexity), Bingbot (Microsoft Copilot), and Google-Extended, which Google documents as governing Gemini grounding as well as training. GPTBot is documented as OpenAI's training crawler and ClaudeBot as feeding model training, so blocking those is a separate decision and does not by itself remove you from ChatGPT or Claude answers. Google-Extended does not affect inclusion or ranking in Google Search. The checker lists each one's status.
Why should I allow AI crawlers in robots.txt?
Blocking what an engine uses to find and cite pages — the crawlers OAI-SearchBot, Claude-SearchBot, PerplexityBot and Bingbot, plus Google-Extended, which is a control token rather than a crawler (Google documents that it has no user-agent string of its own) — keeps your content out of that engine's answers. Blocking a training crawler such as GPTBot or ClaudeBot does not by itself affect citation; it is a separate choice about whether your content trains those models.
Is the robots.txt checker free?
Yes, free and no signup required.