How to Fix AI Crawlers Blocked in robots.txt
Your robots.txt disallows one or more of the crawlers AI answers are drawn from: OAI-SearchBot (ChatGPT), PerplexityBot, Claude-SearchBot, Bingbot, Applebot, DuckAssistBot, Amzn-SearchBot, Meta-WebIndexer or MistralAI-Index. The issue on your report names the exact ones we found blocked. Pages those crawlers cannot fetch cannot be surfaced or linked in the answers they power. The fix is to allow those user-agents in robots.txt, which is a separate decision from whether you allow AI training crawlers - blocking training bots is not what raises this issue.
·
What this means
Robots.txt is a plain-text file at the root of your domain (yoursite.com/robots.txt) that tells automated crawlers which paths they may fetch. Each block starts with a User-agent: line naming a bot, followed by Allow: and Disallow: rules.
This critical fires for one group of bots only: the crawlers AI answers are drawn from. Those are OAI-SearchBot (ChatGPT search), PerplexityBot, Claude-SearchBot, bingbot, Applebot, DuckAssistBot, Amzn-SearchBot, Meta-WebIndexer and MistralAI-Index. Most of them build an index an engine answers from; DuckAssistBot is different - DuckDuckGo says it "crawls pages in real-time for our AI-assisted answers, which prominently cite their sources", and documents no index of its own (DuckDuckGo help pages, read 2026-09-23). Either way, a block means the engine cannot reach your pages. The patterns that trigger it are:
- A named block, such as
User-agent: PerplexityBotfollowed byDisallow: / - A catch-all
User-agent: *withDisallow: /, which applies to each of those bots that has no group of its own in your file - A rule that blocks everything without saying
Disallow: /literally, such asDisallow: /*
We read your file with the same longest-match parser the crawl itself is gated on, and we only call a bot blocked when your homepage is disallowed for it, or when every page we sampled across two or more sections of your site is. A single disallowed folder is never reported as a site-wide block.
Three other kinds of block do not raise this critical, because they do not cost you citations. Each gets its own separate notice on your report instead:
- Training crawlers -
GPTBot,ClaudeBot,CCBot,Amazonbot,Applebot-Extended,meta-externalagent,MistralAI-Training,Bytespider. They collect content for model training. Blocking them is a normal content-licensing choice and does not remove you from AI answers. - Fetchers that run when a person asks -
ChatGPT-User,Perplexity-User,Claude-User,Meta-ExternalFetcher,Amzn-User,Google-Agent,MistralAI-User. These load a single page because someone asked an assistant about it. OpenAI, Perplexity, Meta, Amazon and Google all state that robots.txt may not be enforced for their user-triggered fetchers, so a block here may not even hold. - Google-Extended - it controls Gemini training and grounding in Gemini apps, and has no crawler of its own. Google states it does not impact a site's inclusion in Google Search and is not used as a ranking signal (Google crawler documentation, updated 2026-07-14).
That distinction matters most when something else manages your robots.txt for you. Cloudflare's managed robots.txt adds Disallow: / for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent (Cloudflare documentation, updated 2026-08-03). Every one of those is a training or opt-out token, so that switch on its own does not raise this critical.
One limit worth stating plainly: robots.txt is a voluntary standard. It records your policy, it does not enforce anything at your server. If you need a block enforced, do it at your firewall or CDN.
Why it matters
AI answers link to pages an engine's crawler has already fetched. The crawler named in your report is how your pages reach that engine, and each operator says so in its own documentation: OpenAI recommends allowing OAI-SearchBot if you want to appear in ChatGPT search results; Perplexity and Anthropic run PerplexityBot and Claude-SearchBot for the same job; Microsoft reports where publisher content shows up across Copilot, AI-generated summaries in Bing and partner integrations (Bing Webmaster Tools, 2026-02-10); Apple's Applebot powers Spotlight, Siri and Safari search. A Disallow: / aimed at one of those user-agents is a standing instruction to stay out of that engine's answers.
Being fetchable is a precondition, not a promise. Allowing these crawlers does not guarantee you will be cited; nothing does. What a block does guarantee is that the engine cannot fetch your pages, so it has nothing of yours to surface.
There is a classic-search angle whenever the cause is a catch-all rule. A User-agent: * with Disallow: / does not only stop AI crawlers, it stops Googlebot, which quietly removes you from ordinary search results too. Google's AI Overviews and AI Mode are built on the same index Googlebot populates, so that one rule cuts both channels at once. There is no separate AI Overviews crawler to allow: whether your pages may appear there is a Search Console setting (Settings > Search generative AI) that no crawler can see, so no audit, ours included, can check it for you.
What this issue is not: a verdict on AI training. If you disallow GPTBot, ClaudeBot, CCBot or Google-Extended deliberately, that is a licensing decision, it is reported separately as a notice, and it is not what raised this critical.
How to fix it
- 1
Read your current robots.txt
Open yoursite.com/robots.txt in a browser to see the live file. Look for any
Disallow: /lines and note whichUser-agentblock each sits under. A rule is scoped to the nearestUser-agentline above it, soUser-agent: GPTBotthenDisallow: /blocks only GPTBot, whileUser-agent: *thenDisallow: /blocks everything. Identify whether you have named AI-bot blocks, a global block, or both. - 2
Remove or narrow the blocking rules
For each AI or search crawler you want to allow, either delete its
Disallow: /block or replace/with the specific paths you actually want kept private, like/admin/or/cart/. If a catch-allUser-agent: *withDisallow: /is the culprit, change it toDisallow:(empty, meaning allow all) or scope it to real private directories. An emptyDisallow:and no rule at all both mean fully allowed. - 3
Allow the crawlers this issue is about
This critical clears when the crawlers behind AI answers can fetch your pages. Their tokens are
OAI-SearchBot(OpenAI),PerplexityBot(Perplexity),Claude-SearchBot(Anthropic),bingbot(Microsoft),Applebot(Apple),DuckAssistBot(DuckDuckGo),Amzn-SearchBot(Amazon),Meta-WebIndexer(Meta) andMistralAI-Index(Mistral). Your report names which of them your file actually blocks, and those are the only ones you need to change. KeepGooglebotallowed as well: it is the baseline for Google Search, and for AI Overviews and AI Mode, which run on the same index. The training tokens (GPTBot,ClaudeBot,CCBot,Amazonbot,Applebot-Extended,meta-externalagent,MistralAI-Training,Bytespider) andGoogle-Extendedare a separate choice: leaving those disallowed is normal, and it will not keep this critical on your report. See the code example. - 4
Apply the change on your platform
On WordPress with Yoast or Rank Math, edit robots.txt through the plugin's file editor rather than fighting the virtual file WordPress generates. On Shopify, robots.txt is templated, so edit
robots.txt.liquidin your theme. On Wix and Squarespace, use the built-in robots.txt editor in SEO settings. On raw HTML, Next.js, or any custom stack, edit the physical/public/robots.txt(or your framework's route handler) and redeploy. The file must be reachable at the domain root, not in a subfolder. - 5
Validate and re-run the audit
Re-fetch yoursite.com/robots.txt and confirm the AI user-agents no longer carry
Disallow: /. Watch for ordering traps: within one user-agent block a more specificAllow:can override a broaderDisallow:, but rules across separate blocks do not interact. Run the audit again to confirm the flag clears. Changes take effect whenever each bot next re-reads your robots.txt, which they do on their own schedule (often within a day).
Example
# Allow everything by default, block only genuinely private areas
User-agent: *
Disallow: /admin/
Disallow: /cart/
Disallow: /checkout/
# --- Crawlers behind AI answers ---
# These are the ones this critical fires for.
# OpenAI - ChatGPT search results
User-agent: OAI-SearchBot
Allow: /
# Perplexity
User-agent: PerplexityBot
Allow: /
# Anthropic - Claude search
User-agent: Claude-SearchBot
Allow: /
# Microsoft - the Bing index, and the Copilot answers built on it
User-agent: bingbot
Allow: /
# Apple - Spotlight, Siri, Safari
User-agent: Applebot
Allow: /
# DuckDuckGo AI-assisted answers
User-agent: DuckAssistBot
Allow: /
# Amazon search experiences
User-agent: Amzn-SearchBot
Allow: /
# Meta AI search
User-agent: Meta-WebIndexer
Allow: /
# Mistral search
User-agent: MistralAI-Index
Allow: /
# Classic search - keep this allowed; AI Overviews and AI Mode use the same index
User-agent: Googlebot
Allow: /
# --- Optional: opt out of AI training only ---
# Blocking these does NOT raise the critical above and does not remove
# you from AI answers. Uncomment only if that is your licensing choice.
# User-agent: GPTBot
# Disallow: /
#
# User-agent: ClaudeBot
# Disallow: /
#
# User-agent: CCBot
# Disallow: /
#
# User-agent: Google-Extended
# Disallow: /
# Point crawlers to your sitemap
Sitemap: https://www.yoursite.com/sitemap.xmlA robots.txt that allows the crawlers behind AI answers while keeping private paths blocked. The training section at the bottom is optional: blocking those tokens is a licensing choice, and it is not what raises this issue.
Platform-specific steps
WordPress serves a virtual robots.txt by default. In Yoast, go to Yoast SEO > Tools > File editor to create and edit a real robots.txt. In Rank Math, go to Rank Math > General Settings > Edit robots.txt. Remove any Disallow: / under AI or * user-agents and save. Avoid editing at the server if a plugin is managing the file, or the two will conflict.
Shopify generates robots.txt from a template. In your admin, go to Online Store > Themes > Edit code, then create or open templates/robots.txt.liquid. Add or remove Disallow/Allow rules for the AI user-agents in Liquid, save, and the changes render at yourstore.com/robots.txt.
Both offer a built-in robots.txt editor in SEO settings (Wix: Marketing & SEO > SEO Tools > Robots.txt Editor; Squarespace exposes it under crawling/SEO settings). Edit the AI-crawler rules there. You cannot upload a physical file to the domain root on these hosts, so the built-in editor is the only supported path.
For a static file, edit /public/robots.txt and redeploy. For a dynamic route in Next.js App Router, edit app/robots.ts (or robots.js) so the returned rules allow the AI user-agents, then redeploy. Verify the file resolves at the exact domain root with a 200 status and Content-Type: text/plain.
Frequently asked
That is your call, and it is not what this issue flags. GPTBot and ClaudeBot are training crawlers: disallowing them tells OpenAI and Anthropic your content should not be used to train models, which many publishers choose, and it does not remove you from ChatGPT, Claude or Perplexity answers. The crawlers behind this critical are the ones AI answers are drawn from: OAI-SearchBot, PerplexityBot, Claude-SearchBot, bingbot, Applebot, DuckAssistBot, Amzn-SearchBot, Meta-WebIndexer and MistralAI-Index. If you want to stay citable without being used for training, allow those and disallow the training tokens.
No. Those are training and opt-out tokens. If your report shows them blocked you will see separate notices about AI training crawlers and about Google-Extended, not this critical. Google states Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal (Google crawler documentation, updated 2026-07-14), so it cannot remove you from AI Overviews either. Cloudflare's managed robots.txt switch disallows Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent (Cloudflare documentation, updated 2026-08-03), which are all training or opt-out tokens, so that switch on its own does not raise this critical.
Blocking the dedicated AI tokens (GPTBot, ClaudeBot, Google-Extended) does not affect classic Google rankings, because those are separate from Googlebot. But a broad User-agent: * with Disallow: / blocks Googlebot too, which does destroy rankings. Google-Extended only controls whether your content trains Gemini; it has no effect on Search or on whether you appear in AI Overviews, both of which depend on the regular Googlebot crawl.
Not instantly. Crawlers cache robots.txt and re-fetch it on their own schedules, so a change usually propagates within hours to a day as each bot re-reads the file. Robots.txt is also a voluntary standard. Most major bots honor it, but some have been documented ignoring it or masking their identity, so if you need hard enforcement, block by user-agent or IP at the server or CDN level instead of relying on robots.txt alone.
Both are OpenAI's, with different jobs. GPTBot is the training crawler that gathers content to train future models. OAI-SearchBot indexes pages so they can appear in ChatGPT search results, and it is the one this critical fires for. ChatGPT-User is a third token: it fetches a single page when a ChatGPT user's action needs it, and OpenAI notes that because those fetches are initiated by a person, robots.txt rules may not apply to them. If your goal is to appear in ChatGPT's answers with a link back, OAI-SearchBot is the one to keep allowed. OpenAI documents the three tokens separately so you can control each independently.
At the root of your domain, reachable at exactly yoursite.com/robots.txt. It cannot sit in a subfolder, and it only covers its own host. A subdomain like blog.yoursite.com needs its own separate robots.txt. Confirm the file returns a 200 status as plain text, not a redirect or an HTML error page.
Does your site have this issue?
Run a free, AI-powered audit and we’ll flag this and 270+ other checks. No signup.