Free tool
Is my site blocked from ChatGPT and AI crawlers?
Last updated
Enter a domain. We read its robots.txt the way each AI crawler does, then request the homepage as a browser and as six AI crawlers, and report what came back. A site can allow every crawler in robots.txt and still refuse them at its CDN, so both are checked.
Eight requests: robots.txt, the homepage as a browser, and the homepage as six AI crawlers. Nothing is stored.
What the check measures
- robots.txt. For each crawler, whether the file disallows the site root, and the user-agent group and rule that decide it. A crawler with its own group ignores the
*group, which is the most common reason a site blocks a crawler it meant to allow. - The edge. The homepage, requested once with a normal browser user agent and once with each AI crawler's user agent. A 200 for the browser against a 403, a challenge page or a dropped connection for a crawler is a block that robots.txt does not show, typically a CDN or firewall rule.
Evidence, not proof
The edge probes come from our server with a crawler's user agent, not from the crawler's own network. Some CDNs verify a crawler by its IP address, so the real crawler can be treated differently from our imitation of it, in either direction. A refusal here is strong evidence of a user-agent rule; the CDN's own logs show what the real crawler got.
Googlebot and bingbot are deliberately not probed. Cloudflare and others verify them by IP, so a request that claims to be Googlebot from anywhere else is refused by design and would prove nothing. Their robots.txt verdicts are still reported.
Which crawler does what
- OAI-SearchBot
- OpenAI. Fetches pages for ChatGPT search answers and citations.
- GPTBot
- OpenAI. Collects training data for OpenAI's models.
- Claude-SearchBot
- Anthropic. Indexes pages for Claude's search answers.
- Claude-User
- Anthropic. Fetches a page when a Claude user's question needs it.
- ClaudeBot
- Anthropic. Collects training data for Claude.
- PerplexityBot
- Perplexity. Indexes pages for Perplexity's answers.
- Googlebot
- Google. Its index is what AI Overviews and AI Mode answer from.
- Google-Extended
- Google. A robots.txt token only. It gates grounding in the Gemini apps and has no effect on AI Overviews or AI Mode.
One engine at a time
Each answer engine has its own crawlers and its own index. These pages say which ones decide whether that engine can use a site, from the operator's own documentation, and run the same check with that engine's crawlers first.
Why it matters
An answer engine can only cite a page its crawler could read. A blocked search crawler does not lower the site in that engine's answers; it removes the site from them. Training crawlers are different: blocking them is a legitimate choice that affects what future models learn, not whether today's answers can cite you.
Crawler access is the first of many things that decide AI visibility. The free audit checks the rest on up to 25 pages: whether the content is in the HTML or only rendered by JavaScript, whether sections are short enough to be quoted, and whether pages carry the dates and evidence engines look for. The ai page check shows those measurements for one page.