Which AI crawlers have visited my site?
Last updated
Your server's access log already records every visit from the crawlers behind ChatGPT, Claude, Perplexity, Google's AI Overviews and Copilot. Drop the log here and see which AI crawlers came, on which days, which pages they fetched most, and which status codes they got back. A crawler that receives 403s or 5xx errors never saw the page it asked for.
It is free, there is no signup, and the log is read by this page inside your browser. It is never uploaded.
Drop access log files here
Plain text: nginx, Apache, Cloudflare, Vercel, Netlify or CloudFront. Several files (a rotated log) are read as one.
Read in this tab only. Nothing is uploaded: the page makes no request with your log.
Results appear here: which AI crawlers came, when, what they fetched and what they got back.
What it reports
- Each AI crawler or fetcher found in the log, by its user-agent token, with the operator and what the operator says it is for.
- Requests per crawler per day (UTC), and its share of every request in the log.
- The pages each crawler fetched most, with query strings removed.
- The status codes each crawler received, and the pages that answered it with a 4xx or 5xx error.
- The split of AI requests between training crawlers, search-index crawlers and fetchers a person triggered by asking an assistant.
- The crawlers that did not appear, stated as not seen in this log, with its line count and date range.
Why the status codes matter: an answer engine can only quote what its crawler received. A 403 from a firewall rule or a 503 from an overloaded server means that request got an error page instead of the content, whatever the page says when you open it yourself.
Supported log formats
Plain-text files, one record per line. Several files are read as one, so a rotated access.log and access.log.1 can go in together. Compressed files (.gz, .zip, .zst) are recognised and refused with a message: decompress them first.
- nginx, Apache
- The combined log format, as both servers predefine it, including a virtual-host prefix, IPv6 addresses, escaped quotes and extra fields after the user agent (nginx, Apache). The common format has no user agent, so it cannot be used.
- Cloudflare
- Logpush HTTP requests as JSON lines, with ClientRequestUserAgent, ClientRequestPath or ClientRequestURI, EdgeResponseStatus and EdgeStartTimestamp in any timestamp format (field reference).
- Vercel
- Log Drain output, JSON or NDJSON; request entries carry a proxy object with the path, user agent and status. Build and function output lines are skipped (schema).
- Netlify
- Log Drain traffic records, JSON or NDJSON, with url, status_code and user_agent. A drain set to exclude personal data omits the user agent and cannot be used (log drains).
- Amazon CloudFront, IIS
- W3C extended logs with a #Fields header, tab- or space-separated; URL-encoded user agents are decoded (CloudFront standard logs).
Your log stays in your browser
The file is read by JavaScript on this page, line by line, and only the counts are kept. No part of the log is sent to GetCited or anyone else: the analyzer makes no network request at all. To check, open your browser's network panel before you drop a file, or disconnect from the internet once the page has loaded; it works the same.
This site measures its own pages with a cookieless behaviour snippet that records clicks and copied text. On this page it is kept away from the paste box and the results, so it cannot record any of your log. Large files are streamed rather than loaded whole; the largest measured so far was a 500 MB nginx log of 2.3 million lines, read in about 7 seconds in desktop Chrome with the page responsive throughout.
The AI crawlers it recognises
Each token below was checked against the operator's own documentation on 4 October 2026. A request is matched on the whole token, so Googlebot-Image is not counted as Googlebot.
| Token | Operator | Purpose | What the operator says |
|---|---|---|---|
| GPTBot | OpenAI | Training | Crawls content that may be used to train OpenAI's foundation models. Source |
| OAI-SearchBot | OpenAI | Search index | Surfaces websites in ChatGPT's search features. Source |
| ChatGPT-User | OpenAI | User-triggered | Visits a page for a user action in ChatGPT or a custom GPT; not an automatic crawl. Source |
| ClaudeBot | Anthropic | Training | Collects web content that could contribute to training Anthropic's models. Source |
| Claude-SearchBot | Anthropic | Search index | Navigates the web to improve the quality of Claude's search results. Source |
| Claude-User | Anthropic | User-triggered | Visits a page when a person asks Claude a question that needs it. Source |
| PerplexityBot | Perplexity | Search index | Surfaces and links websites in Perplexity's search results. Source |
| Perplexity-User | Perplexity | User-triggered | Visits a page to answer a question a person asked in Perplexity. Source |
| Googlebot | Search index | Google Search's crawler. AI Overviews and AI Mode answer from the index it builds. Source | |
| Google-Extended | robots.txt only | robots.txt control over Gemini model training and Gemini app grounding. Requests use Google's existing user agents. Source | |
| bingbot | Microsoft | Search index | Bing's crawler. Bing's index is what Copilot answers retrieve from. Source |
| Applebot | Apple | Search index | Powers search in Spotlight, Siri and Safari. Source |
| Applebot-Extended | Apple | robots.txt only | robots.txt control over whether pages Applebot crawled train Apple's foundation models. It does not crawl. Source |
| Meta-ExternalAgent | Meta | Training | Crawls for uses such as training foundation AI models or improving products. Source |
| Meta-WebIndexer | Meta | Search index | Navigates the web to improve Meta AI search results. Source |
| Meta-ExternalFetcher | Meta | User-triggered | Fetches individual links at a user's request. Source |
| Amazonbot | Amazon | Training | Improves Amazon products and services; may be used to train Amazon AI models. Source |
| Amzn-SearchBot | Amazon | Search index | Search experiences in Amazon products; not used for model training. Source |
| Amzn-User | Amazon | User-triggered | Supports user actions, such as Alexa answers that need current information. Source |
| DuckAssistBot | DuckDuckGo | Search index | Fetches sources in real time for DuckDuckGo's AI-assisted answers; not used to train models. Source |
| MistralAI-Training | Mistral AI | Training | Builds datasets for training Mistral's generative models. Source |
| MistralAI-Index | Mistral AI | Search index | Crawls for indexing only, to support Mistral's search. Source |
| MistralAI-User | Mistral AI | User-triggered | Visits a page a user asked for while using Mistral's assistant. Source |
| CCBot | Common Crawl | Training | Builds Common Crawl's open web archive, a common source of model training data. Source |
| Bytespider | ByteDance | Training | ByteDance's crawler, widely reported to collect training data for its models. Third-party source (ByteDance publishes no crawler documentation we could reach on 2026-10-04; the linked page is a third-party description.) |
What a log can and cannot tell you
- A user agent can be forged. It is whatever the client chose to send. A request is proven to come from an operator only when its IP address falls in the ranges that operator publishes; the results link to them where they exist. This tool does not check IP addresses.
- Not seen is not blocked. A crawler missing from the log made no request that reached the server that wrote it. A CDN or firewall in front of that server may have answered it first, and only the CDN's own log would show that.
- Some tokens never appear in a log. Google-Extended and Applebot-Extended are robots.txt controls: Google and Apple crawl with their usual user agents and apply the control afterwards.
- A fetch is not a citation. A visit from a search or user-triggered fetcher shows the page was retrieved, not that an answer quoted it.
What a log cannot measure
The log shows that a crawler asked for a page and what status it got. It cannot show whether the page was usable once fetched: whether its text is in the HTML without JavaScript, whether its title, headings and dates say what it is, whether robots.txt and the page itself let answer engines use it. GetCited's free first audit fetches your site the way an answer engine does and reports exactly that.