Free tool
AI crawler directory
Last verified
The crawlers and fetchers that AI companies send to websites, 29 of them from 12 operators: the token robots.txt rules address, the user agent that shows up in your logs, what the operator says the requests are for, and where it publishes the addresses they come from.
Every row was checked against the operator's own documentation, linked in the row, on the date shown. Nothing is taken from third-party lists or from traffic logs. Where an operator does not say something, the row says “Not documented by the operator” rather than guess.
To see which of these a particular site's robots.txt lets in, use the AI crawler checker.
To check whether an address in your logs really belongs to the crawler it claims to be, use Is this really GPTBot?, which tests it against these published ranges.
All crawlers at a glance
| Crawler | Used for | robots.txt |
|---|---|---|
| GPTBotOpenAI | Model training | Yes |
| OAI-SearchBotOpenAI | Search index | Yes |
| ChatGPT-UserOpenAI | User-triggered fetch | Not always |
| OAI-AdsBotOpenAI | Other | Not documented |
| ClaudeBotAnthropic | Model training | Yes |
| Claude-SearchBotAnthropic | Search index | Yes |
| Claude-UserAnthropic | User-triggered fetch | Yes |
| PerplexityBotPerplexity | Search index | Yes |
| Perplexity-UserPerplexity | User-triggered fetch | No |
| GooglebotGoogle | Search index | Yes |
| Google-ExtendedGoogle | Model training, Grounding | Yes |
| Google-CloudVertexBotGoogle | Other | Yes |
| Google-AgentGoogle | User-triggered fetch | No |
| Google-GeminiNotebookGoogle | User-triggered fetch | No |
| bingbotMicrosoft | Search index | Yes |
| ApplebotApple | Search index, Model training, Grounding | Yes |
| Applebot-ExtendedApple | Model training | Yes |
| meta-externalagentMeta | Model training, Other | Yes |
| meta-webindexerMeta | Search index | Yes |
| meta-externalfetcherMeta | User-triggered fetch | Not always |
| AmazonbotAmazon | Model training, Other | Yes |
| Amzn-SearchBotAmazon | Search index | Yes |
| Amzn-UserAmazon | User-triggered fetch | Not always |
| CCBotCommon Crawl | Open dataset | Yes |
| MistralAI-UserMistral | User-triggered fetch | Yes |
| MistralAI-IndexMistral | Search index | Not documented |
| MistralAI-TrainingMistral | Model training | Yes |
| DuckAssistBotDuckDuckGo | Search index | Yes |
| YouBotYou.com | Search index | Yes |
How to read the table
- Model training: the operator says content it collects may train its models. Search index: it builds the index an AI answer or search product retrieves from. User-triggered fetch: it requests a page only when a person asks something. Grounding: it decides whether content may back a model's answers. Open dataset: a public crawl others build on.
- Several operators run separate tokens for training and for search, and say each one is independent: OpenAI (GPTBot, OAI-SearchBot), Anthropic (ClaudeBot, Claude-SearchBot), Mistral (MistralAI-Training, MistralAI-Index) and Amazon (Amazonbot, Amzn-SearchBot).
- Google-Extended is not a crawler. It governs Gemini training and grounding in Gemini Apps and on Vertex AI. It has no effect on AI Overviews or AI Mode, which use Googlebot's index. Applebot-Extended is likewise a control token for Apple's model training, not a crawler.
- User-triggered fetchers often do not follow robots.txt, by the operators' own account, because a person asked for the page. Each row quotes what its operator says.
- “Runs JavaScript” is filled in only where the operator says so. Most AI operators do not document it.
OpenAI
GPTBot
OpenAI's training crawler. It crawls content that may be used to train OpenAI's generative AI foundation models. Disallowing GPTBot indicates a site's content should not be used for that training. OpenAI treats it as independent of OAI-SearchBot.
- Operator
- OpenAI
- Kind
- Crawler: fetches on its own schedule
- Used for
- Model training
- Feeds
- OpenAI foundation model training
- robots.txt token
GPTBot- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
- Follows robots.txt
- Yes. OpenAI documents the GPTBot robots.txt tag as the training control.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- gptbot.json: https://openai.com/gptbot.jsonCheck an address against these ranges
- Official source
- Overview of OpenAI Crawlershttps://developers.openai.com/api/docs/bots
- Verified
OAI-SearchBot
OpenAI's search crawler. It is used to surface websites in ChatGPT's search features. OpenAI says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers, though they can still appear as navigational links. Robots.txt changes take about 24 hours to apply.
- Operator
- OpenAI
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- ChatGPT search
- robots.txt token
OAI-SearchBot- User agent string
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
- Follows robots.txt
- Yes. OpenAI documents the OAI-SearchBot robots.txt tag as the search control.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- searchbot.json: https://openai.com/searchbot.jsonCheck an address against these ranges
- Official source
- Overview of OpenAI Crawlershttps://developers.openai.com/api/docs/bots
- Verified
ChatGPT-User
Fetches a page when a user asks ChatGPT or a Custom GPT something, or uses a GPT Action. OpenAI says it is not used to crawl the web automatically and not used to decide whether content may appear in ChatGPT search; that is OAI-SearchBot.
- Operator
- OpenAI
- Kind
- Fetcher: requests a page when a user asks
- Used for
- User-triggered fetch
- Feeds
- ChatGPT; Custom GPTs; GPT Actions
- robots.txt token
ChatGPT-User- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
- Follows robots.txt
- Not always. OpenAI: because these actions are initiated by a user, robots.txt rules may not apply.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- chatgpt-user.json: https://openai.com/chatgpt-user.jsonCheck an address against these ranges
- Official source
- Overview of OpenAI Crawlershttps://developers.openai.com/api/docs/bots
- Verified
OAI-AdsBot
Visits the landing pages of ads submitted to ChatGPT, to check them against OpenAI's policies and to judge when an ad is relevant. OpenAI says it only visits pages submitted as ads and its data is not used to train foundation models.
- Operator
- OpenAI
- Kind
- Crawler: fetches on its own schedule
- Used for
- Other
- Feeds
- Ads in ChatGPT
- robots.txt token
- Not documented by the operator. It is identified by its user agent only.
- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot
- Follows robots.txt
- Not documented by the operator
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- adsbot.json: https://openai.com/adsbot.jsonCheck an address against these ranges
- Official source
- Overview of OpenAI Crawlershttps://developers.openai.com/api/docs/bots
- Verified
Anthropic
ClaudeBot
Anthropic's training crawler. It collects web content that could contribute to training Anthropic's generative AI models. Anthropic says restricting ClaudeBot signals that the site's future materials should be excluded from its training datasets.
- Operator
- Anthropic
- Kind
- Crawler: fetches on its own schedule
- Used for
- Model training
- Feeds
- Anthropic model training
- robots.txt token
ClaudeBot- User agent string
- Anthropic names the robots.txt token but does not publish a full user agent string.
- Follows robots.txt
- Yes. Anthropic says its bots honour robots.txt directives and the non-standard Crawl-delay extension.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- One list for all of Anthropic's bots: https://claude.com/crawling/bots.jsonCheck an address against these ranges
- Official source
- Does Anthropic crawl data from the web, and how can site owners block the crawler?https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Verified
Claude-SearchBot
Anthropic's search crawler. It navigates the web to improve the relevance and accuracy of search responses in Claude. Anthropic says disabling it prevents indexing the site for search, which may reduce its visibility in user search results.
- Operator
- Anthropic
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- Claude search results
- robots.txt token
Claude-SearchBot- User agent string
- Anthropic names the robots.txt token but does not publish a full user agent string.
- Follows robots.txt
- Yes. Anthropic says its bots honour robots.txt directives and the non-standard Crawl-delay extension.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- One list for all of Anthropic's bots: https://claude.com/crawling/bots.jsonCheck an address against these ranges
- Official source
- Does Anthropic crawl data from the web, and how can site owners block the crawler?https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Verified
Claude-User
Fetches pages when a person asks Claude a question. Anthropic says disabling it stops Claude retrieving the site's content in response to a user query, which may reduce the site's visibility for user-directed web search.
- Operator
- Anthropic
- Kind
- Fetcher: requests a page when a user asks
- Used for
- User-triggered fetch
- Feeds
- Claude
- robots.txt token
Claude-User- User agent string
- Anthropic names the robots.txt token but does not publish a full user agent string.
- Follows robots.txt
- Yes. Anthropic says its bots, including this one, honour robots.txt directives.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- One list for all of Anthropic's bots: https://claude.com/crawling/bots.jsonCheck an address against these ranges
- Official source
- Does Anthropic crawl data from the web, and how can site owners block the crawler?https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Verified
Perplexity
PerplexityBot
Perplexity's search crawler. It surfaces and links websites in Perplexity's search results. Perplexity says it is not used to crawl content for AI foundation models. Robots.txt changes can take up to 24 hours to apply.
- Operator
- Perplexity
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- Perplexity search results
- robots.txt token
PerplexityBot- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
- Follows robots.txt
- Yes. Perplexity documents the PerplexityBot robots.txt tag as its control.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- perplexitybot.json: https://www.perplexity.com/perplexitybot.jsonCheck an address against these ranges
- Official source
- Perplexity Crawlershttps://docs.perplexity.ai/docs/resources/perplexity-crawlers
- Verified
Perplexity-User
Fetches a page when a user asks Perplexity a question, to answer it and link the page in the response. Perplexity says it is not used for web crawling or to collect content for training AI foundation models.
- Operator
- Perplexity
- Kind
- Fetcher: requests a page when a user asks
- Used for
- User-triggered fetch
- Feeds
- Perplexity answers
- robots.txt token
Perplexity-User- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
- Follows robots.txt
- No. Perplexity: since a user requested the fetch, this fetcher generally ignores robots.txt rules.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- perplexity-user.json: https://www.perplexity.com/perplexity-user.jsonCheck an address against these ranges
- Official source
- Perplexity Crawlershttps://docs.perplexity.ai/docs/resources/perplexity-crawlers
- Verified
Googlebot
Google Search's crawler, and the crawler behind Google's AI features in Search. Google says robots.txt rules for Googlebot are the control for how a site is crawled for Search, AI Overviews and AI Mode included, and a page must be indexed and eligible for a snippet to be a supporting link there.
- Operator
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- Google Search; AI Overviews; AI Mode; Discover; Google Images; Google News
- robots.txt token
Googlebot- User agent strings
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36
- Follows robots.txt
- Yes. Google: common crawlers always obey robots.txt when crawling automatically.
- Runs JavaScript
- Yes. Google: a headless Chromium renders the page and executes the JavaScript, after a queue.
- Published IP ranges
- common-crawlers.json: https://developers.google.com/static/crawling/ipranges/common-crawlers.jsonCheck an address against these ranges
- Official sources
- List of Google's common crawlershttps://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
- AI features and your websitehttps://developers.google.com/search/docs/appearance/ai-features
- Understand JavaScript SEO basicshttps://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics
- Verified
Google-Extended
A robots.txt control token, not a crawler. It governs whether content Google crawls may train future Gemini models and ground answers in Gemini Apps and Grounding with Google Search on Vertex AI. It has no effect on Google Search, AI Overviews or AI Mode, which use Googlebot's index.
- Operator
- Kind
- Control token: read in robots.txt, sends no requests
- Used for
- Model training, Grounding
- Feeds
- Gemini Apps (training and grounding); Vertex AI API for Gemini (training); Grounding with Google Search on Vertex AI
- robots.txt token
Google-Extended- User agent string
- No user agent of its own. Google crawls with its existing user agents and reads this token only as a control in robots.txt.
- Follows robots.txt
- Yes. Google documents the token for use in robots.txt only.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- None of its own: the crawling is done by Google's common crawlers: https://developers.google.com/static/crawling/ipranges/common-crawlers.json
- Official sources
- List of Google's common crawlershttps://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
- AI features and your websitehttps://developers.google.com/search/docs/appearance/ai-features
- Verified
Google-CloudVertexBot
Crawls that a site's owner has requested for building Vertex AI Agents on Google Cloud. Google says it has no effect on Google Search or other products.
- Operator
- Kind
- Crawler: fetches on its own schedule
- Used for
- Other
- Feeds
- Vertex AI Agents
- robots.txt token
Google-CloudVertexBot- User agent string
- Google publishes only the substring Google-CloudVertexBot, not a full user agent string.
- Follows robots.txt
- Yes. Google: common crawlers always obey robots.txt when crawling automatically.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- common-crawlers.json: https://developers.google.com/static/crawling/ipranges/common-crawlers.jsonCheck an address against these ranges
- Official source
- List of Google's common crawlershttps://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers
- Verified
Google-Agent
Used by agents hosted on Google infrastructure to navigate the web and perform actions when a user asks. Google lists it among its user-triggered fetchers and is also experimenting with Web Bot Auth for it, under the identity https://agent.bot.goog.
- Operator
- Kind
- Fetcher: requests a page when a user asks
- Used for
- User-triggered fetch
- Feeds
- Agents hosted on Google infrastructure
- robots.txt token
- Not documented by the operator. It is identified by its user agent only.
- User agent strings
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent)Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko; compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent) Chrome/W.X.Y.Z Safari/537.36
- Follows robots.txt
- No. Google: because the fetch was requested by a user, user-triggered fetchers generally ignore robots.txt rules.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- user-triggered-agents.json: https://developers.google.com/static/crawling/ipranges/user-triggered-agents.jsonCheck an address against these ranges
- Official source
- List of Google user-triggered fetchershttps://developers.google.com/crawling/docs/crawlers-fetchers/google-user-triggered-fetchers
- Verified
Google-GeminiNotebook
Requests the individual URLs that Gemini Notebook users add as sources for their projects. Google lists the former agent name, Google-NotebookLM, as supported until August 2026.
- Operator
- Kind
- Fetcher: requests a page when a user asks
- Used for
- User-triggered fetch
- Feeds
- Gemini Notebook
- robots.txt token
- Not documented by the operator. It is identified by its user agent only.
- User agent strings
Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/138.0.0.0 Mobile Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)
- Follows robots.txt
- No. Google: because the fetch was requested by a user, user-triggered fetchers generally ignore robots.txt rules.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- user-triggered-fetchers-google.json (Google lists three user-triggered files and does not say which this fetcher uses): https://developers.google.com/static/crawling/ipranges/user-triggered-fetchers-google.jsonCheck an address against these ranges
- Official source
- List of Google user-triggered fetchershttps://developers.google.com/crawling/docs/crawlers-fetchers/google-user-triggered-fetchers
- Verified
Microsoft
bingbot
Bing's standard crawler, which handles most of its crawling. Microsoft says pages removed from the Bing index stop appearing in Bing results and in Copilot experiences that rely on Bing's index.
- Operator
- Microsoft
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- Bing search; Copilot experiences that rely on Bing's index
- robots.txt token
bingbot- User agent strings
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)
- Follows robots.txt
- Yes. Bing: robots.txt can be configured to tell Bing crawlers how to interact with a site.
- Runs JavaScript
- Yes. Bing says it keeps its page rendering engine on the latest stable Microsoft Edge.
- Published IP ranges
- bingbot.json: https://www.bing.com/toolbox/bingbot.jsonCheck an address against these ranges
- Official sources
- Which Crawlers Does Bing Use?https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0
- How To Permanently Remove a URL or Page from Bing or Copilothttps://www.bing.com/webmasters/help/how-to-permanently-remove-a-url-or-page-from-bing-or-copilot-37c07477
- How to Verify Bingbothttps://www.bing.com/webmasters/help/how-to-verify-bingbot-3905dc26
- Verified
Apple
Applebot
Apple's crawler. Its data powers search in Spotlight, Siri and Safari, may help train Apple's foundation models, and may give AI models context for answers in Siri and Search. Training is controlled separately through Applebot-Extended.
- Operator
- Apple
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index, Model training, Grounding
- Feeds
- Spotlight; Siri; Safari; AI-generated answers in Siri and Search; Apple foundation model training (opt-out token: Applebot-Extended)
- robots.txt token
Applebot- User agent strings
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)Mozilla/5.0 (iPhone; CPU iPhone OS 17_4_1 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4.1 Mobile/15E148 Safari/604.1 (Applebot/0.1; +http://www.apple.com/go/applebot)
- Follows robots.txt
- Yes. Apple: respects robots.txt rules aimed at Applebot, and follows Googlebot's rules when Applebot is not named.
- Runs JavaScript
- Yes. Apple: Applebot may render the content of your website within a browser.
- Published IP ranges
- applebot.json: https://search.developer.apple.com/applebot.jsonCheck an address against these ranges
- Official source
- About Applebothttps://support.apple.com/en-us/119829
- Verified
Applebot-Extended
A robots.txt control token, not a crawler. It governs whether content Applebot crawls may train Apple's foundation models for Apple Intelligence, Services and Developer Tools. Apple says pages that disallow it can still appear in search results.
- Operator
- Apple
- Kind
- Control token: read in robots.txt, sends no requests
- Used for
- Model training
- Feeds
- Apple foundation model training
- robots.txt token
Applebot-Extended- User agent string
- Apple: Applebot-Extended does not crawl webpages. It is a robots.txt control token only.
- Follows robots.txt
- Yes. Apple documents the token for use in robots.txt only.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- None of its own: the crawling is done by Applebot: https://search.developer.apple.com/applebot.json
- Official source
- About Applebothttps://support.apple.com/en-us/119829
- Verified
Meta
meta-externalagent
Meta's crawler for use cases such as training foundation AI models or improving products by indexing content directly. Meta says it may cache robots.txt for up to 24 hours.
- Operator
- Meta
- Kind
- Crawler: fetches on its own schedule
- Used for
- Model training, Other
- Feeds
- Meta foundation model training; Meta products (indexing)
- robots.txt token
meta-externalagent- User agent strings
meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)meta-externalagent/1.1
- Follows robots.txt
- Yes. Meta: add a disallow for the relevant crawler to block it.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- Not documented by the operator
- Official source
- Meta Web Crawlershttps://developers.facebook.com/documentation/sharing/webmasters/web-crawlers
- Verified
meta-webindexer
Navigates the web to improve the quality of Meta AI search results. Meta says allowing it helps Meta cite and link to the site's content in Meta AI's responses.
- Operator
- Meta
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- Meta AI
- robots.txt token
meta-webindexer- User agent strings
meta-webindexer/1.1 (+/documentation/sharing/webmasters/web-crawlers)meta-webindexer/1.1
- Follows robots.txt
- Yes. Meta: add a disallow for the relevant crawler to block it.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- Not documented by the operator
- Official source
- Meta Web Crawlershttps://developers.facebook.com/documentation/sharing/webmasters/web-crawlers
- Verified
meta-externalfetcher
Fetches individual links at a user's request and supports product functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users.
- Operator
- Meta
- Kind
- Fetcher: requests a page when a user asks
- Used for
- User-triggered fetch
- Feeds
- Meta AI agentic features
- robots.txt token
meta-externalfetcher- User agent strings
meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers)meta-externalfetcher/1.1
- Follows robots.txt
- Not always. Meta: may bypass robots.txt because it performs fetches that were requested by the user.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- Not documented by the operator
- Official source
- Meta Web Crawlershttps://developers.facebook.com/documentation/sharing/webmasters/web-crawlers
- Verified
Amazon
Amazonbot
Used to improve Amazon's products and services, which Amazon says helps it give customers more accurate information and may be used to train Amazon AI models. Amazon says it honours the noarchive meta tag as do not use for model training.
- Operator
- Amazon
- Kind
- Crawler: fetches on its own schedule
- Used for
- Model training, Other
- Feeds
- Amazon products and services; Amazon AI model training
- robots.txt token
Amazonbot- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36
- Follows robots.txt
- Yes. Amazon: automated crawling respects the Robots Exclusion Protocol; crawl-delay is not supported.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- Amazonbot IP addresses: https://developer.amazon.com/amazonbot/ip-addresses/
- Official source
- Amazonbothttps://developer.amazon.com/amazonbot
- Verified
Amzn-SearchBot
Used to improve search experiences in Amazon products and services, such as Alexa. Amazon says it does not crawl content for generative AI model training, and follows the rules given to other search bots when robots.txt does not name it.
- Operator
- Amazon
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- Alexa; Amazon search experiences
- robots.txt token
Amzn-SearchBot- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36
- Follows robots.txt
- Yes. Amazon: automated crawling respects the Robots Exclusion Protocol; crawl-delay is not supported.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- Amzn-SearchBot IP addresses: https://developer.amazon.com/amazonbot/searchbot-ip-addresses/
- Official source
- Amazonbothttps://developer.amazon.com/amazonbot
- Verified
Amzn-User
Supports user actions, such as fetching live information from the web to answer an Alexa question on the user's behalf. Amazon says it does not crawl content for generative AI model training.
- Operator
- Amazon
- Kind
- Fetcher: requests a page when a user asks
- Used for
- User-triggered fetch
- Feeds
- Alexa
- robots.txt token
Amzn-User- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36
- Follows robots.txt
- Not always. Amazon: because its actions can be initiated by a user, it may not follow all robots.txt directives.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- Amzn-User IP addresses: https://developer.amazon.com/amazonbot/live-ip-addresses/
- Official source
- Amazonbothttps://developer.amazon.com/amazonbot
- Verified
Common Crawl
CCBot
The crawler of Common Crawl, a non-profit that maintains an open repository of web crawl data anyone can access and analyse. Others build on that repository; Common Crawl's CCBot page does not name the uses.
- Operator
- Common Crawl
- Kind
- Crawler: fetches on its own schedule
- Used for
- Open dataset
- Feeds
- Common Crawl open repository
- robots.txt token
CCBot- User agent string
CCBot/2.0 (https://commoncrawl.org/faq/)
- Follows robots.txt
- Yes. Common Crawl documents a CCBot robots.txt group to prevent crawling.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- ccbot.json (IPv4 and IPv6): https://index.commoncrawl.org/ccbot.jsonCheck an address against these ranges
- Official source
- CCBothttps://commoncrawl.org/ccbot
- Verified
Mistral
MistralAI-User
Visits a page when a user asks Mistral's Vibe a question, and may link the source in the answer. Mistral says it is not used to crawl the web automatically nor to collect content for generative AI training.
- Operator
- Mistral
- Kind
- Fetcher: requests a page when a user asks
- Used for
- User-triggered fetch
- Feeds
- Vibe
- robots.txt token
MistralAI-User- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)
- Follows robots.txt
- Yes. Mistral: MistralAI-User governs which sites these user requests can be made to.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- mistralai-user-ips.json: https://mistral.ai/mistralai-user-ips.jsonCheck an address against these ranges
- Official source
- Mistral crawlershttps://docs.mistral.ai/robots
- Verified
MistralAI-Index
Crawls the web automatically, for indexing only, to build Mistral search, which helps answer questions in Vibe. Mistral says content it crawls is not used for generative AI training of any kind.
- Operator
- Mistral
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- Mistral search; Vibe
- robots.txt token
MistralAI-Index- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)
- Follows robots.txt
- Not documented by the operator
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- mistralai-index-ips.json: https://mistral.ai/mistralai-index-ips.jsonCheck an address against these ranges
- Official source
- Mistral crawlershttps://docs.mistral.ai/robots
- Verified
MistralAI-Training
Crawls web content to help build datasets for training Mistral's generative AI models. Mistral says it is not used for search indexing or to answer live questions in Vibe.
- Operator
- Mistral
- Kind
- Crawler: fetches on its own schedule
- Used for
- Model training
- Feeds
- Mistral model training
- robots.txt token
MistralAI-Training- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)
- Follows robots.txt
- Yes. Mistral: webmasters can disallow this user agent in robots.txt.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- Not documented by the operator
- Official source
- Mistral crawlershttps://docs.mistral.ai/robots
- Verified
DuckDuckGo
DuckAssistBot
Crawls pages in real time for DuckDuckGo Search's AI-assisted answers, which cite their sources. DuckDuckGo says the data is not used to train AI models, and opting out does not affect organic rankings.
- Operator
- DuckDuckGo
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- DuckDuckGo AI-assisted answers
- robots.txt token
DuckAssistBot- User agent string
DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)
- Follows robots.txt
- Yes. DuckDuckGo: a robots.txt disallow takes effect after 72 hours.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- duckassistbot.json: https://duckduckgo.com/duckassistbot.jsonCheck an address against these ranges
- Official source
- Is DuckAssistBot related to DuckDuckGo?https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot
- Verified
You.com
YouBot
The crawler that powers the You.com search engine, discovering and indexing pages for its results. You.com says it caches robots.txt for 30 minutes and signs its requests with Web Bot Auth.
- Operator
- You.com
- Kind
- Crawler: fetches on its own schedule
- Used for
- Search index
- Feeds
- You.com search
- robots.txt token
YouBot- User agent string
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; YouBot/1.0; +https://docs.you.com/youbot; env:prod) Chrome/X.X.X.X Safari/537.36
- Follows robots.txt
- Yes. You.com: respects robots.txt, including user-agent rules and crawl-delay.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- 68.67.112.0/24, stated on the documentation pageCheck an address against these ranges
- Official source
- YouBot: You.com's Web Crawlerhttps://docs.you.com/youbot
- Verified
Crawlers not listed
A crawler is listed only when its operator documents it. These are often asked about and are left out for that reason:
Bytespider: Attributed to ByteDance, but we could not find crawler documentation published by ByteDance itself, so nothing about it could be verified at the source.cohere-ai: Attributed to Cohere, but we could not find crawler documentation published by Cohere itself.
How this list is kept
Each row stores the operator's page it was checked against and the date of the check. The whole table was last verified on . Operators rename and add agents; a row older than its operator's page may be out of date, and the operator's page is the authority.
What this page does not do
It does not say which crawlers to allow or block. That depends on what a site is for, and the operators' pages describe what each choice changes.
Your site's AI visibility
Whether an answer engine can use a page once its crawler has it is a separate question from access. Your first GetCited audit is free: one site, up to 25 pages, every finding with its evidence, ready for your coding agent.