AI crawler directory
What is CCBot?
Last verified
The crawler of Common Crawl, a non-profit that maintains an open repository of web crawl data anyone can access and analyse. Others build on that repository; Common Crawl's CCBot page does not name the uses.
CCBot at a glance
- Operator
- Common Crawl
- Kind
- Crawler: fetches on its own schedule
- Used for
- Open dataset
- Feeds
- Common Crawl open repository
- robots.txt token
CCBot- User agent string
CCBot/2.0 (https://commoncrawl.org/faq/)
- Follows robots.txt
- Yes. Common Crawl documents a CCBot robots.txt group to prevent crawling.
- Runs JavaScript
- Not documented by the operator
- Published IP ranges
- ccbot.json (IPv4 and IPv6): https://index.commoncrawl.org/ccbot.jsonCheck an address against these ranges
- Official source
- CCBothttps://commoncrawl.org/ccbot
- Verified
Every AI crawler in one table
CCBot is one of 29 crawlers and fetchers in the AI crawler directory, each checked against its operator's own documentation.