GETCITED

Free tool

AI crawler directory

Last verified

The crawlers and fetchers that AI companies send to websites, 29 of them from 12 operators: the token robots.txt rules address, the user agent that shows up in your logs, what the operator says the requests are for, and where it publishes the addresses they come from.

Every row was checked against the operator's own documentation, linked in the row, on the date shown. Nothing is taken from third-party lists or from traffic logs. Where an operator does not say something, the row says “Not documented by the operator” rather than guess.

To see which of these a particular site's robots.txt lets in, use the AI crawler checker.

To check whether an address in your logs really belongs to the crawler it claims to be, use Is this really GPTBot?, which tests it against these published ranges.

All crawlers at a glance

AI crawlers by operator, what each is used for, and whether it follows robots.txt
CrawlerUsed forrobots.txt
GPTBotOpenAIModel trainingYes
OAI-SearchBotOpenAISearch indexYes
ChatGPT-UserOpenAIUser-triggered fetchNot always
OAI-AdsBotOpenAIOtherNot documented
ClaudeBotAnthropicModel trainingYes
Claude-SearchBotAnthropicSearch indexYes
Claude-UserAnthropicUser-triggered fetchYes
PerplexityBotPerplexitySearch indexYes
Perplexity-UserPerplexityUser-triggered fetchNo
GooglebotGoogleSearch indexYes
Google-ExtendedGoogleModel training, GroundingYes
Google-CloudVertexBotGoogleOtherYes
Google-AgentGoogleUser-triggered fetchNo
Google-GeminiNotebookGoogleUser-triggered fetchNo
bingbotMicrosoftSearch indexYes
ApplebotAppleSearch index, Model training, GroundingYes
Applebot-ExtendedAppleModel trainingYes
meta-externalagentMetaModel training, OtherYes
meta-webindexerMetaSearch indexYes
meta-externalfetcherMetaUser-triggered fetchNot always
AmazonbotAmazonModel training, OtherYes
Amzn-SearchBotAmazonSearch indexYes
Amzn-UserAmazonUser-triggered fetchNot always
CCBotCommon CrawlOpen datasetYes
MistralAI-UserMistralUser-triggered fetchYes
MistralAI-IndexMistralSearch indexNot documented
MistralAI-TrainingMistralModel trainingYes
DuckAssistBotDuckDuckGoSearch indexYes
YouBotYou.comSearch indexYes

How to read the table

  • Model training: the operator says content it collects may train its models. Search index: it builds the index an AI answer or search product retrieves from. User-triggered fetch: it requests a page only when a person asks something. Grounding: it decides whether content may back a model's answers. Open dataset: a public crawl others build on.
  • Several operators run separate tokens for training and for search, and say each one is independent: OpenAI (GPTBot, OAI-SearchBot), Anthropic (ClaudeBot, Claude-SearchBot), Mistral (MistralAI-Training, MistralAI-Index) and Amazon (Amazonbot, Amzn-SearchBot).
  • Google-Extended is not a crawler. It governs Gemini training and grounding in Gemini Apps and on Vertex AI. It has no effect on AI Overviews or AI Mode, which use Googlebot's index. Applebot-Extended is likewise a control token for Apple's model training, not a crawler.
  • User-triggered fetchers often do not follow robots.txt, by the operators' own account, because a person asked for the page. Each row quotes what its operator says.
  • “Runs JavaScript” is filled in only where the operator says so. Most AI operators do not document it.

OpenAI

GPTBot

OpenAI's training crawler. It crawls content that may be used to train OpenAI's generative AI foundation models. Disallowing GPTBot indicates a site's content should not be used for that training. OpenAI treats it as independent of OAI-SearchBot.

Operator
OpenAI
Kind
Crawler: fetches on its own schedule
Used for
Model training
Feeds
OpenAI foundation model training
robots.txt token
GPTBot
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
Follows robots.txt
Yes. OpenAI documents the GPTBot robots.txt tag as the training control.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is GPTBot? Its own page

OAI-SearchBot

OpenAI's search crawler. It is used to surface websites in ChatGPT's search features. OpenAI says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers, though they can still appear as navigational links. Robots.txt changes take about 24 hours to apply.

Operator
OpenAI
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
ChatGPT search
robots.txt token
OAI-SearchBot
User agent string
  • Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
Follows robots.txt
Yes. OpenAI documents the OAI-SearchBot robots.txt tag as the search control.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is OAI-SearchBot? Its own page

ChatGPT-User

Fetches a page when a user asks ChatGPT or a Custom GPT something, or uses a GPT Action. OpenAI says it is not used to crawl the web automatically and not used to decide whether content may appear in ChatGPT search; that is OAI-SearchBot.

Operator
OpenAI
Kind
Fetcher: requests a page when a user asks
Used for
User-triggered fetch
Feeds
ChatGPT; Custom GPTs; GPT Actions
robots.txt token
ChatGPT-User
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
Follows robots.txt
Not always. OpenAI: because these actions are initiated by a user, robots.txt rules may not apply.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is ChatGPT-User? Its own page

OAI-AdsBot

Visits the landing pages of ads submitted to ChatGPT, to check them against OpenAI's policies and to judge when an ad is relevant. OpenAI says it only visits pages submitted as ads and its data is not used to train foundation models.

Operator
OpenAI
Kind
Crawler: fetches on its own schedule
Used for
Other
Feeds
Ads in ChatGPT
robots.txt token
Not documented by the operator. It is identified by its user agent only.
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot
Follows robots.txt
Not documented by the operator
Runs JavaScript
Not documented by the operator
Official source
Verified

What is OAI-AdsBot? Its own page

Anthropic

ClaudeBot

Anthropic's training crawler. It collects web content that could contribute to training Anthropic's generative AI models. Anthropic says restricting ClaudeBot signals that the site's future materials should be excluded from its training datasets.

Operator
Anthropic
Kind
Crawler: fetches on its own schedule
Used for
Model training
Feeds
Anthropic model training
robots.txt token
ClaudeBot
User agent string
Anthropic names the robots.txt token but does not publish a full user agent string.
Follows robots.txt
Yes. Anthropic says its bots honour robots.txt directives and the non-standard Crawl-delay extension.
Runs JavaScript
Not documented by the operator
Published IP ranges
One list for all of Anthropic's bots: https://claude.com/crawling/bots.jsonCheck an address against these ranges
Official source
Verified

What is ClaudeBot? Its own page

Claude-SearchBot

Anthropic's search crawler. It navigates the web to improve the relevance and accuracy of search responses in Claude. Anthropic says disabling it prevents indexing the site for search, which may reduce its visibility in user search results.

Operator
Anthropic
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
Claude search results
robots.txt token
Claude-SearchBot
User agent string
Anthropic names the robots.txt token but does not publish a full user agent string.
Follows robots.txt
Yes. Anthropic says its bots honour robots.txt directives and the non-standard Crawl-delay extension.
Runs JavaScript
Not documented by the operator
Published IP ranges
One list for all of Anthropic's bots: https://claude.com/crawling/bots.jsonCheck an address against these ranges
Official source
Verified

What is Claude-SearchBot? Its own page

Claude-User

Fetches pages when a person asks Claude a question. Anthropic says disabling it stops Claude retrieving the site's content in response to a user query, which may reduce the site's visibility for user-directed web search.

Operator
Anthropic
Kind
Fetcher: requests a page when a user asks
Used for
User-triggered fetch
Feeds
Claude
robots.txt token
Claude-User
User agent string
Anthropic names the robots.txt token but does not publish a full user agent string.
Follows robots.txt
Yes. Anthropic says its bots, including this one, honour robots.txt directives.
Runs JavaScript
Not documented by the operator
Published IP ranges
One list for all of Anthropic's bots: https://claude.com/crawling/bots.jsonCheck an address against these ranges
Official source
Verified

What is Claude-User? Its own page

Perplexity

PerplexityBot

Perplexity's search crawler. It surfaces and links websites in Perplexity's search results. Perplexity says it is not used to crawl content for AI foundation models. Robots.txt changes can take up to 24 hours to apply.

Operator
Perplexity
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
Perplexity search results
robots.txt token
PerplexityBot
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Follows robots.txt
Yes. Perplexity documents the PerplexityBot robots.txt tag as its control.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is PerplexityBot? Its own page

Perplexity-User

Fetches a page when a user asks Perplexity a question, to answer it and link the page in the response. Perplexity says it is not used for web crawling or to collect content for training AI foundation models.

Operator
Perplexity
Kind
Fetcher: requests a page when a user asks
Used for
User-triggered fetch
Feeds
Perplexity answers
robots.txt token
Perplexity-User
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
Follows robots.txt
No. Perplexity: since a user requested the fetch, this fetcher generally ignores robots.txt rules.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is Perplexity-User? Its own page

Google

Googlebot

Google Search's crawler, and the crawler behind Google's AI features in Search. Google says robots.txt rules for Googlebot are the control for how a site is crawled for Search, AI Overviews and AI Mode included, and a page must be indexed and eligible for a snippet to be a supporting link there.

Operator
Google
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
Google Search; AI Overviews; AI Mode; Discover; Google Images; Google News
robots.txt token
Googlebot
User agent strings
  • Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36
Follows robots.txt
Yes. Google: common crawlers always obey robots.txt when crawling automatically.
Runs JavaScript
Yes. Google: a headless Chromium renders the page and executes the JavaScript, after a queue.
Official sources
Verified

What is Googlebot? Its own page

Google-Extended

A robots.txt control token, not a crawler. It governs whether content Google crawls may train future Gemini models and ground answers in Gemini Apps and Grounding with Google Search on Vertex AI. It has no effect on Google Search, AI Overviews or AI Mode, which use Googlebot's index.

Operator
Google
Kind
Control token: read in robots.txt, sends no requests
Used for
Model training, Grounding
Feeds
Gemini Apps (training and grounding); Vertex AI API for Gemini (training); Grounding with Google Search on Vertex AI
robots.txt token
Google-Extended
User agent string
No user agent of its own. Google crawls with its existing user agents and reads this token only as a control in robots.txt.
Follows robots.txt
Yes. Google documents the token for use in robots.txt only.
Runs JavaScript
Not documented by the operator
Published IP ranges
None of its own: the crawling is done by Google's common crawlers: https://developers.google.com/static/crawling/ipranges/common-crawlers.json
Official sources
Verified

What is Google-Extended? Its own page

Google-CloudVertexBot

Crawls that a site's owner has requested for building Vertex AI Agents on Google Cloud. Google says it has no effect on Google Search or other products.

Operator
Google
Kind
Crawler: fetches on its own schedule
Used for
Other
Feeds
Vertex AI Agents
robots.txt token
Google-CloudVertexBot
User agent string
Google publishes only the substring Google-CloudVertexBot, not a full user agent string.
Follows robots.txt
Yes. Google: common crawlers always obey robots.txt when crawling automatically.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is Google-CloudVertexBot? Its own page

Google-Agent

Used by agents hosted on Google infrastructure to navigate the web and perform actions when a user asks. Google lists it among its user-triggered fetchers and is also experimenting with Web Bot Auth for it, under the identity https://agent.bot.goog.

Operator
Google
Kind
Fetcher: requests a page when a user asks
Used for
User-triggered fetch
Feeds
Agents hosted on Google infrastructure
robots.txt token
Not documented by the operator. It is identified by its user agent only.
User agent strings
  • Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent)
  • Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko; compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent) Chrome/W.X.Y.Z Safari/537.36
Follows robots.txt
No. Google: because the fetch was requested by a user, user-triggered fetchers generally ignore robots.txt rules.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is Google-Agent? Its own page

Google-GeminiNotebook

Requests the individual URLs that Gemini Notebook users add as sources for their projects. Google lists the former agent name, Google-NotebookLM, as supported until August 2026.

Operator
Google
Kind
Fetcher: requests a page when a user asks
Used for
User-triggered fetch
Feeds
Gemini Notebook
robots.txt token
Not documented by the operator. It is identified by its user agent only.
User agent strings
  • Mozilla/5.0 (Linux; Android 10; K) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/138.0.0.0 Mobile Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)
  • Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)
Follows robots.txt
No. Google: because the fetch was requested by a user, user-triggered fetchers generally ignore robots.txt rules.
Runs JavaScript
Not documented by the operator
Published IP ranges
user-triggered-fetchers-google.json (Google lists three user-triggered files and does not say which this fetcher uses): https://developers.google.com/static/crawling/ipranges/user-triggered-fetchers-google.jsonCheck an address against these ranges
Official source
Verified

What is Google-GeminiNotebook? Its own page

Microsoft

bingbot

Bing's standard crawler, which handles most of its crawling. Microsoft says pages removed from the Bing index stop appearing in Bing results and in Copilot experiences that rely on Bing's index.

Operator
Microsoft
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
Bing search; Copilot experiences that rely on Bing's index
robots.txt token
bingbot
User agent strings
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/W.X.Y.Z Safari/537.36
  • Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)
  • Mozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)
Follows robots.txt
Yes. Bing: robots.txt can be configured to tell Bing crawlers how to interact with a site.
Runs JavaScript
Yes. Bing says it keeps its page rendering engine on the latest stable Microsoft Edge.
Official sources
Verified

What is bingbot? Its own page

Apple

Applebot

Apple's crawler. Its data powers search in Spotlight, Siri and Safari, may help train Apple's foundation models, and may give AI models context for answers in Siri and Search. Training is controlled separately through Applebot-Extended.

Operator
Apple
Kind
Crawler: fetches on its own schedule
Used for
Search index, Model training, Grounding
Feeds
Spotlight; Siri; Safari; AI-generated answers in Siri and Search; Apple foundation model training (opt-out token: Applebot-Extended)
robots.txt token
Applebot
User agent strings
  • Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)
  • Mozilla/5.0 (iPhone; CPU iPhone OS 17_4_1 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4.1 Mobile/15E148 Safari/604.1 (Applebot/0.1; +http://www.apple.com/go/applebot)
Follows robots.txt
Yes. Apple: respects robots.txt rules aimed at Applebot, and follows Googlebot's rules when Applebot is not named.
Runs JavaScript
Yes. Apple: Applebot may render the content of your website within a browser.
Official source
Verified

What is Applebot? Its own page

Applebot-Extended

A robots.txt control token, not a crawler. It governs whether content Applebot crawls may train Apple's foundation models for Apple Intelligence, Services and Developer Tools. Apple says pages that disallow it can still appear in search results.

Operator
Apple
Kind
Control token: read in robots.txt, sends no requests
Used for
Model training
Feeds
Apple foundation model training
robots.txt token
Applebot-Extended
User agent string
Apple: Applebot-Extended does not crawl webpages. It is a robots.txt control token only.
Follows robots.txt
Yes. Apple documents the token for use in robots.txt only.
Runs JavaScript
Not documented by the operator
Published IP ranges
None of its own: the crawling is done by Applebot: https://search.developer.apple.com/applebot.json
Official source
Verified

What is Applebot-Extended? Its own page

Meta

meta-externalagent

Meta's crawler for use cases such as training foundation AI models or improving products by indexing content directly. Meta says it may cache robots.txt for up to 24 hours.

Operator
Meta
Kind
Crawler: fetches on its own schedule
Used for
Model training, Other
Feeds
Meta foundation model training; Meta products (indexing)
robots.txt token
meta-externalagent
User agent strings
  • meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)
  • meta-externalagent/1.1
Follows robots.txt
Yes. Meta: add a disallow for the relevant crawler to block it.
Runs JavaScript
Not documented by the operator
Published IP ranges
Not documented by the operator
Official source
  • Meta Web Crawlershttps://developers.facebook.com/documentation/sharing/webmasters/web-crawlers
Verified

What is meta-externalagent? Its own page

meta-webindexer

Navigates the web to improve the quality of Meta AI search results. Meta says allowing it helps Meta cite and link to the site's content in Meta AI's responses.

Operator
Meta
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
Meta AI
robots.txt token
meta-webindexer
User agent strings
  • meta-webindexer/1.1 (+/documentation/sharing/webmasters/web-crawlers)
  • meta-webindexer/1.1
Follows robots.txt
Yes. Meta: add a disallow for the relevant crawler to block it.
Runs JavaScript
Not documented by the operator
Published IP ranges
Not documented by the operator
Official source
  • Meta Web Crawlershttps://developers.facebook.com/documentation/sharing/webmasters/web-crawlers
Verified

What is meta-webindexer? Its own page

meta-externalfetcher

Fetches individual links at a user's request and supports product functions such as evaluating and improving agentic AI capabilities, including helping AI navigate websites to complete tasks for users.

Operator
Meta
Kind
Fetcher: requests a page when a user asks
Used for
User-triggered fetch
Feeds
Meta AI agentic features
robots.txt token
meta-externalfetcher
User agent strings
  • meta-externalfetcher/1.1 (+/documentation/sharing/webmasters/web-crawlers)
  • meta-externalfetcher/1.1
Follows robots.txt
Not always. Meta: may bypass robots.txt because it performs fetches that were requested by the user.
Runs JavaScript
Not documented by the operator
Published IP ranges
Not documented by the operator
Official source
  • Meta Web Crawlershttps://developers.facebook.com/documentation/sharing/webmasters/web-crawlers
Verified

What is meta-externalfetcher? Its own page

Amazon

Amazonbot

Used to improve Amazon's products and services, which Amazon says helps it give customers more accurate information and may be used to train Amazon AI models. Amazon says it honours the noarchive meta tag as do not use for model training.

Operator
Amazon
Kind
Crawler: fetches on its own schedule
Used for
Model training, Other
Feeds
Amazon products and services; Amazon AI model training
robots.txt token
Amazonbot
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36
Follows robots.txt
Yes. Amazon: automated crawling respects the Robots Exclusion Protocol; crawl-delay is not supported.
Runs JavaScript
Not documented by the operator
Published IP ranges
Amazonbot IP addresses: https://developer.amazon.com/amazonbot/ip-addresses/
Official source
  • Amazonbothttps://developer.amazon.com/amazonbot
Verified

What is Amazonbot? Its own page

Amzn-SearchBot

Used to improve search experiences in Amazon products and services, such as Alexa. Amazon says it does not crawl content for generative AI model training, and follows the rules given to other search bots when robots.txt does not name it.

Operator
Amazon
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
Alexa; Amazon search experiences
robots.txt token
Amzn-SearchBot
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36
Follows robots.txt
Yes. Amazon: automated crawling respects the Robots Exclusion Protocol; crawl-delay is not supported.
Runs JavaScript
Not documented by the operator
Published IP ranges
Amzn-SearchBot IP addresses: https://developer.amazon.com/amazonbot/searchbot-ip-addresses/
Official source
  • Amazonbothttps://developer.amazon.com/amazonbot
Verified

What is Amzn-SearchBot? Its own page

Amzn-User

Supports user actions, such as fetching live information from the web to answer an Alexa question on the user's behalf. Amazon says it does not crawl content for generative AI model training.

Operator
Amazon
Kind
Fetcher: requests a page when a user asks
Used for
User-triggered fetch
Feeds
Alexa
robots.txt token
Amzn-User
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36
Follows robots.txt
Not always. Amazon: because its actions can be initiated by a user, it may not follow all robots.txt directives.
Runs JavaScript
Not documented by the operator
Published IP ranges
Amzn-User IP addresses: https://developer.amazon.com/amazonbot/live-ip-addresses/
Official source
  • Amazonbothttps://developer.amazon.com/amazonbot
Verified

What is Amzn-User? Its own page

Common Crawl

CCBot

The crawler of Common Crawl, a non-profit that maintains an open repository of web crawl data anyone can access and analyse. Others build on that repository; Common Crawl's CCBot page does not name the uses.

Operator
Common Crawl
Kind
Crawler: fetches on its own schedule
Used for
Open dataset
Feeds
Common Crawl open repository
robots.txt token
CCBot
User agent string
  • CCBot/2.0 (https://commoncrawl.org/faq/)
Follows robots.txt
Yes. Common Crawl documents a CCBot robots.txt group to prevent crawling.
Runs JavaScript
Not documented by the operator
Official source
  • CCBothttps://commoncrawl.org/ccbot
Verified

What is CCBot? Its own page

Mistral

MistralAI-User

Visits a page when a user asks Mistral's Vibe a question, and may link the source in the answer. Mistral says it is not used to crawl the web automatically nor to collect content for generative AI training.

Operator
Mistral
Kind
Fetcher: requests a page when a user asks
Used for
User-triggered fetch
Feeds
Vibe
robots.txt token
MistralAI-User
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)
Follows robots.txt
Yes. Mistral: MistralAI-User governs which sites these user requests can be made to.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is MistralAI-User? Its own page

MistralAI-Index

Crawls the web automatically, for indexing only, to build Mistral search, which helps answer questions in Vibe. Mistral says content it crawls is not used for generative AI training of any kind.

Operator
Mistral
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
Mistral search; Vibe
robots.txt token
MistralAI-Index
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)
Follows robots.txt
Not documented by the operator
Runs JavaScript
Not documented by the operator
Official source
Verified

What is MistralAI-Index? Its own page

MistralAI-Training

Crawls web content to help build datasets for training Mistral's generative AI models. Mistral says it is not used for search indexing or to answer live questions in Vibe.

Operator
Mistral
Kind
Crawler: fetches on its own schedule
Used for
Model training
Feeds
Mistral model training
robots.txt token
MistralAI-Training
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)
Follows robots.txt
Yes. Mistral: webmasters can disallow this user agent in robots.txt.
Runs JavaScript
Not documented by the operator
Published IP ranges
Not documented by the operator
Official source
Verified

What is MistralAI-Training? Its own page

DuckDuckGo

DuckAssistBot

Crawls pages in real time for DuckDuckGo Search's AI-assisted answers, which cite their sources. DuckDuckGo says the data is not used to train AI models, and opting out does not affect organic rankings.

Operator
DuckDuckGo
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
DuckDuckGo AI-assisted answers
robots.txt token
DuckAssistBot
User agent string
  • DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)
Follows robots.txt
Yes. DuckDuckGo: a robots.txt disallow takes effect after 72 hours.
Runs JavaScript
Not documented by the operator
Official source
Verified

What is DuckAssistBot? Its own page

You.com

YouBot

The crawler that powers the You.com search engine, discovering and indexing pages for its results. You.com says it caches robots.txt for 30 minutes and signs its requests with Web Bot Auth.

Operator
You.com
Kind
Crawler: fetches on its own schedule
Used for
Search index
Feeds
You.com search
robots.txt token
YouBot
User agent string
  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; YouBot/1.0; +https://docs.you.com/youbot; env:prod) Chrome/X.X.X.X Safari/537.36
Follows robots.txt
Yes. You.com: respects robots.txt, including user-agent rules and crawl-delay.
Runs JavaScript
Not documented by the operator
Published IP ranges
68.67.112.0/24, stated on the documentation pageCheck an address against these ranges
Official source
Verified

What is YouBot? Its own page

Crawlers not listed

A crawler is listed only when its operator documents it. These are often asked about and are left out for that reason:

  • Bytespider: Attributed to ByteDance, but we could not find crawler documentation published by ByteDance itself, so nothing about it could be verified at the source.
  • cohere-ai: Attributed to Cohere, but we could not find crawler documentation published by Cohere itself.

How this list is kept

Each row stores the operator's page it was checked against and the date of the check. The whole table was last verified on . Operators rename and add agents; a row older than its operator's page may be out of date, and the operator's page is the authority.

What this page does not do

It does not say which crawlers to allow or block. That depends on what a site is for, and the operators' pages describe what each choice changes.

Your site's AI visibility

Whether an answer engine can use a page once its crawler has it is a separate question from access. Your first GetCited audit is free: one site, up to 25 pages, every finding with its evidence, ready for your coding agent.

Audit your AI visibility.

Run the free audit