GETCITED

Free tool

Is this really GPTBot?

Last verified

Paste an IP address, or a whole line from your access log, and see whether that address belongs to the AI crawler it claims to be. The check uses what each operator publishes: its list of IP ranges, and for Google, Microsoft, Apple, Common Crawl and You.com, the reverse DNS name with a forward lookup that has to lead back to the same address.

A user agent is a line of text any client can send. An address in the operator's own published ranges is the evidence that a request came from its crawler.

Check an address

One address (IPv4 or IPv6) or one line in combined log format; the first line is read.

The answer appears here. Nothing you paste is stored.

Why it matters for AI visibility

  • Scrapers routinely send a user agent such as GPTBot or ClaudeBot to borrow a known crawler's reputation. Counting them as real makes AI crawler traffic in your logs look larger, or different, than it is.
  • Some CDNs and firewalls treat a crawler as verified only when its address is in the operator's published ranges. A real crawler from a range the operator added recently, or an impostor with the right user agent, can be treated differently from what the user agent suggests.
  • When an answer engine never cites a page, one question is whether its crawler reached the page at all. Knowing which requests in your logs are genuine tells you which visits to count.

What each crawler is checked against

Generated from the AI crawler directory. A file link is the operator's own range list; a reverse DNS row names the domain the operator says its crawler's hostnames end in.

How each AI crawler in the directory can be verified
CrawlerHow it is checked
GPTBotOpenAIPublished ranges: gptbot.json
OAI-SearchBotOpenAIPublished ranges: searchbot.json
ChatGPT-UserOpenAIPublished ranges: chatgpt-user.json
OAI-AdsBotOpenAIPublished ranges: adsbot.json
ClaudeBotAnthropicPublished ranges: bots.json
Claude-SearchBotAnthropicPublished ranges: bots.json
Claude-UserAnthropicPublished ranges: bots.json
PerplexityBotPerplexityPublished ranges: perplexitybot.json
Perplexity-UserPerplexityPublished ranges: perplexity-user.json
GooglebotGooglePublished ranges: common-crawlers.jsonReverse DNS under googlebot.com (method)
Google-ExtendedGoogleControl token: sends no requests, nothing to verify
Google-CloudVertexBotGooglePublished ranges: common-crawlers.jsonReverse DNS under googlebot.com (method)
Google-AgentGooglePublished ranges: user-triggered-agents.jsonReverse DNS under google.com (method)
Google-GeminiNotebookGooglePublished ranges: user-triggered-fetchers-google.jsonReverse DNS under google.com (method)
bingbotMicrosoftPublished ranges: bingbot.jsonReverse DNS under search.msn.com (method)
ApplebotApplePublished ranges: applebot.jsonReverse DNS under applebot.apple.com (method)
Applebot-ExtendedAppleControl token: sends no requests, nothing to verify
meta-externalagentMetaMeta publishes no IP ranges
meta-webindexerMetaMeta publishes no IP ranges
meta-externalfetcherMetaMeta publishes no IP ranges
AmazonbotAmazonAddresses on a web page only, not a file this tool reads
Amzn-SearchBotAmazonAddresses on a web page only, not a file this tool reads
Amzn-UserAmazonAddresses on a web page only, not a file this tool reads
CCBotCommon CrawlPublished ranges: ccbot.jsonReverse DNS under crawl.commoncrawl.org (method)
MistralAI-UserMistralPublished ranges: mistralai-user-ips.json
MistralAI-IndexMistralPublished ranges: mistralai-index-ips.json
MistralAI-TrainingMistralMistral publishes no IP ranges
DuckAssistBotDuckDuckGoPublished ranges: duckassistbot.json
YouBotYou.comRange stated on its documentationReverse DNS under search.you.com (method)

How the check works

  • The address is read strictly: IPv4 or IPv6, with an IPv4-mapped IPv6 address read as the IPv4 address it carries. Private, loopback, documentation and other reserved addresses are answered as not public, since no crawler reaches the open web from them.
  • The user agent is matched against the crawler tokens in the directory. It is what the request claims; the address is what is checked.
  • Range lists are downloaded by our server from the operators' fixed URLs, at most once every 12 hours each, never once per check. If an operator's file cannot be fetched, the last good copy is used and the answer says when it was fetched. The lists bundled with this page were fetched on .
  • Reverse DNS: the address's hostname is looked up, kept only if it ends in the operator's documented domain, and then resolved forward; it confirms only if the same address comes back. These are DNS lookups only. The tool never connects to the address you paste.

What it cannot tell you

  • Meta and several others publish no IP ranges, and Amazon lists its addresses on web pages rather than in a file. For those, the answer says the claim cannot be verified this way, which is not the same as false.
  • A list is only as current as the operator keeps it. An address the operator started using after its last update reads as not in the ranges.
  • It checks one address at a time and does not read your logs. Nothing you paste is stored or logged.
  • It does not say what to allow or block. That depends on what your site is for; the operators' pages describe what each crawler does.

Your site's AI visibility

Knowing a crawler is real is the first half; whether an answer engine can use what it fetched is the second. Your first GetCited audit is free: one site, up to 25 pages, every finding with its evidence, ready for your coding agent.

Audit your AI visibility.

Run the free audit