Free AI Crawler Log Analyzer
Upload your server access log — or a Cloudflare export — to see which AI crawlers actually visited (GPTBot/OAI-SearchBot, Claude, Perplexity, Google & more), which pages an assistant fetched while answering someone, and how many people arrived from an AI answer. Includes IP verification where the operator publishes usable ranges — we cover OpenAI’s, Perplexity’s and Anthropic’s, refreshed daily. Runs entirely in your browser.
100% private. Your log is parsed entirely in your browser — nothing is uploaded to any server, and no visitor IP is ever shown or stored.
Reads the Combined Log Format that nginx and Apache write, and Cloudflare’s log export — one JSON record per line, which is what Logpush writes by default. Get logs from your host’s file manager, or export them from Cloudflare.
Detects 29+ AI crawlers including GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Bingbot, CCBot & more — plus answer-time fetches and visits referred by ChatGPT, Perplexity, Gemini, Microsoft Copilot and Claude.
How to read your AI crawler report
Each row is one AI crawler the tool matched by its User-Agent, with two counts that mean different things.Hits is every request that bot made; pages is how many distinct URLs it fetched. A bot with many hits but few pages is re-fetching a handful of URLs; a bot with pages close to hits is exploring broadly.Last seen tells you whether the crawl is ongoing or went quiet weeks ago, and the verificationcolumn separates real visits from impostors for the three operators whose range files we bundle: an OpenAI, Perplexity or Anthropic hit is marked verified only if the request IP falls inside the file that operator publishes. Every other bot is matched by User-Agent alone and shown as unverified — that means we did not check, not that the hit is fake. Those three files are re-fetched from the operators daily and the page downloads that copy, with the snapshot built into the page as the fallback. The same word appears for every bot if the copy the tool ends up with ever passes 30 days old: verification switches itself off rather than let an ageing copy call real traffic forged.
Training crawlers and answer engines are not the same visit
The report tags every bot by purpose, and the distinction changes what you should do about it:
- Search crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot, Bingbot, DuckAssistBot) are the crawlers AI answers are drawn from — you want these crawling regularly.
- Training crawlers (GPTBot, CCBot, Bytespider) pull content to train models; blocking them is a legitimate choice that does not cost you citations.
- User-triggered fetches (ChatGPT-User, Claude-User, Perplexity-User) fire when an assistant opens your page while answering someone. Encouraging, but being fetched is not the same as being cited — the assistant may read the page and never link it.
Pages fetched while an assistant was answering someone
Those user-triggered fetches get their own table, page by page, because they are the most interesting thing a log can show you. A scheduled crawl tells you a bot can reach a page. A fetch from ChatGPT-User, Claude-User or Perplexity-User tells you something narrower and better: a person asked a question, and the assistant went and opened that specific page while composing its reply. OpenAI describes ChatGPT-User visiting a page when “users ask ChatGPT or a CustomGPT a question” (developers.openai.com, read 2026-09-23); Perplexity says the same of Perplexity-User (docs.perplexity.ai, read 2026-09-23). Read the list as demand: the pages that keep appearing are the ones people’s questions keep landing on. What it is not is proof of a citation — the answer may have quoted a competitor and never linked you at all.
Visits from AI assistants are people, not crawlers
The last table counts something different again: humans who clicked through to you from an answer, identified by a referrer of chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com or claude.ai — or by a utm_source=chatgpt.com tag on the landing URL when the referrer was stripped. These never share a total with crawler hits, and you should not add them yourself: one is a bot reading a page, the other is a reader arriving. Two caveats come with the number. It is a minimum, because some assistants and in-app browsers send no referrer at all. And it cannot include Google’s AI Overviews or AI Mode, which arrive looking like ordinary Google search — Google Analytics 4 draws the same line, defining its AI Assistant channel as arrivals from sources like ChatGPT, Gemini, Deepseek, Copilot or Grok while excluding AI Overviews and AI Mode (support.google.com, read 2026-09-23). Google says the same on its own side: pages appearing in AI features are reported in Search Console’s Performance report “within the ‘Web’ search type” (developers.google.com, read 2026-09-23), so nobody — including us — can split them out. Whether an answer actually cited you is a different question again, and only Bing publishes that per page: our AI Citation Report reads the Bing Webmaster export for it.
What your most-crawled pages tell you
The most AI-crawled pages list shows where crawlers are spending their budget. If they hammer your homepage but never reach the pages you actually want cited, those pages may be buried, slow, or thin — strengthen internal links to them and keep the content substantive. If a search crawler is missing from the report entirely, the common causes are a robots.txt rule, a firewall or CDN rule, or simply a log that is too short or too old — check the first with our AI Bot Access Checker before assuming anything.
Which log files this reads
Two formats go in: the Combined Log Format that nginx and Apache write by default, and Cloudflare’s log export, which puts one JSON record on each line — Cloudflare documents Logpush emitting “each record as a single line of JSON” (developers.cloudflare.com, read 2026-09-23). The tool prints which of the two it read your file as, so a format it could not parse shows up as a format problem rather than as a quiet week with no crawlers. One thing is worth checking before you export: if your log format is configurable, keep the referrer field. Without it the crawler tables still work, but visits from AI assistants cannot be seen at all, and the tool will say so rather than report zero.
Frequently asked questions
How can I tell if AI crawlers are visiting my website?
AI crawlers identify themselves in the User-Agent header of every request — for example GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot. Your web server records every request (with its User-Agent and IP) in an access log. This tool reads that log and surfaces exactly which AI bots hit your site, how many times, and which pages they fetched.
Why can’t I just use Google Analytics to track AI bots?
Google Analytics runs as a client-side JavaScript tag. OpenAI’s and Anthropic’s crawlers, and Perplexity’s, do not run JavaScript, so it never fires for them — they are invisible to it. The only reliable way to detect bot visits is server-side: your access logs (or your CDN/Cloudflare logs), which is exactly what this tool analyzes.
What does “pages AI assistants fetched while answering someone” mean?
Three fetchers run only when a person asks an assistant a question, rather than on a crawl schedule: ChatGPT-User, Claude-User and Perplexity-User. OpenAI documents ChatGPT-User visiting a web page when “users ask ChatGPT or a CustomGPT a question” (developers.openai.com/api/docs/bots, read 2026-09-23), and Perplexity says Perplexity-User fetches a page “when users ask Perplexity a question” (docs.perplexity.ai/guides/bots, read 2026-09-23). Seeing those hits per page is the closest a server log gets to evidence that your page was read while an answer was being written. It is not proof of a citation: an assistant can read a page and never link it.
Can I see how many visitors arrived from ChatGPT or Perplexity?
Yes, and they are reported separately from crawler hits, because being crawled and being visited are different events that must never be added together. The tool counts requests whose referrer is chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com or claude.ai, plus landing URLs carrying utm_source=chatgpt.com. Treat the result as a minimum: some assistants and in-app browsers send no referrer at all. Google’s AI Overviews and AI Mode cannot be separated this way — Google Analytics 4 draws the same line, defining its AI Assistant channel as arrivals “from sources like ChatGPT, Gemini, Deepseek, Copilot, or Grok” while excluding AI Overviews and AI Mode (support.google.com/analytics/answer/9756891, read 2026-09-23).
Does this work with a Cloudflare log export?
Yes. Cloudflare Logpush writes one JSON record per line by default — “each record as a single line of JSON (also known as ndjson)” (developers.cloudflare.com/logs/logpush/logpush-job/log-output-options/, read 2026-09-23) — and this tool reads that format as well as the Combined Log Format nginx and Apache write. It uses Cloudflare’s own field names, so ClientRequestUserAgent, ClientRequestReferer, ClientRequestURI, ClientIP and EdgeStartTimestamp all come through. The tool tells you which format it read your file as, so a mis-detected export can never look like a quiet week.
What does “verified” vs “spoofed” mean?
Anyone can send a fake User-Agent claiming to be GPTBot. To catch impostors, this tool checks each request’s IP against the range file the operator publishes for that bot. We cover OpenAI’s (GPTBot, OAI-SearchBot, ChatGPT-User), Perplexity’s (PerplexityBot, Perplexity-User) and Anthropic’s (ClaudeBot, Claude-User, Claude-SearchBot): our server re-fetches those files from the operators every day, and the page downloads that copy when it loads, falling back to the snapshot built into the page if the download fails. Everything else shows as “unverified” — that means we did not check, not that the bot is suspect. Google and Microsoft publish files too, but a miss against either cannot support an accusation: Google says its crawlers “generally” crawl from its published ranges (developers.google.com/search/docs/crawling-indexing/google-common-crawlers, read 2026-09-23), and Microsoft’s bingbot.json was stamped January 2024 when we read it on 2026-09-23. Both can still be verified by reverse DNS, which a browser-only tool cannot perform. Whichever copy the tool ends up with, it is used for 30 days and no longer: if our refresh stops, verification switches itself off and everything reads “unverified”, because these ranges churn — about half of OpenAI’s ChatGPT-User prefixes turned over in the 93 days before our 2026-09-23 snapshot — and a stale copy would accuse genuine traffic of being fake. Two innocent explanations exist for an out-of-range hit even so: the operator may have added addresses since our last refresh, or your log may record a proxy’s address rather than the caller’s.
Is my log file uploaded anywhere?
No. The entire analysis runs in your browser using JavaScript — your log file (which may contain visitor IPs) never leaves your device and is never sent to any server. The page does fetch one thing from us, the public list of crawler IP ranges, and that is a download only: no part of your log, and no count taken from it, is ever sent back. No visitor IP is displayed or stored: addresses are used only to compare against a crawler’s published range, and the referral report keeps the landing path and the assistant, never the visitor.
Where do I get my access log?
Most hosts expose it in cPanel/Plesk under “Raw Access Logs,” or at a path like /var/log/nginx/access.log on a VPS. If you’re on Cloudflare, export request logs with Logpush and paste or upload the resulting file — its one-JSON-record-per-line format is read directly. Either way, include the referrer field if your log format is configurable: without it, visits from AI assistants cannot be seen at all.
What should I do if a key AI bot never appears?
A search crawler that never fetches your pages cannot surface them, so it is worth chasing. Check the plain causes in order: how long a window your log actually covers, then your robots.txt (our AI Bot Access Checker reads it per crawler), then any firewall or CDN bot rule — that last one is invisible in robots.txt, and Cloudflare’s free AI Crawl Control shows it for Cloudflare sites. Crawl frequency also varies a lot; absence from one week of logs is not proof of a block.
Related free tools
Instant SEO Snapshot
One-click SEO & AI-readiness grade for any page — HTTPS, indexability, title, meta, headings, content & more. Embeddable on your own site.
AI Bot Access Checker
What does your robots.txt say to OAI-SearchBot, PerplexityBot, Claude-SearchBot, GPTBot and the rest? Read with the real rules.
AI Model Recall Checker
Ask free AI models the unbranded questions your buyers type, and count how many name your brand unprompted — with each model and its training cutoff shown. It measures what those models remember from training, not a live check of ChatGPT, Claude or Perplexity; a built-in panel lets you spot-check the real engines by hand.
NLP Content Analyzer
Topics, entities, search intent & content gaps an NLP model sees in your page.
AI Retrieval Checker
AI-retrieval readiness for one page — an AI model’s read of how findable and quotable it is, with fixes.
AI Citation Report (Bing Copilot)
Upload your Bing Webmaster AI Performance export to see which of your pages Microsoft Copilot and Bing’s AI answers cite, and how that trends. Covers Copilot and Bing only — Google’s AI numbers are in Search Console’s own Generative AI report. Runs in your browser.
Want the full picture?
This tool checks one thing. Run a complete, free SEO audit across 28 modules.
Run a free SEO audit