AI Search Visibility / Crawler Access
Is GPTBot, ClaudeBot or PerplexityBot Blocked From Your Site?
Many sites inherit robots.txt rules for AI-related crawlers without checking what each control actually covers. Paste your URL to see how the site treats 10 named crawlers, with live retrieval, user-triggered access, training and other purposes reported separately.
The Honest Frame
Blocked ≠ invisible to that AI, in every case
Access is necessary but not sufficient, and the reverse isn't a clean rule either. GPTBot feeds OpenAI model training plus some retrieval — blocking it does not switch off every way ChatGPT might reference your site, since browsing/search fetches can use a different user-agent some sites don't think to block. CCBot seeds Common Crawl, which several LLMs train on indirectly, so its effect is diffuse and delayed rather than immediate. This tool reports the one thing that IS a clean, verifiable fact — what your robots.txt actually says about each crawler, quoting the exact line — and explains what that specific rule does and doesn't control, instead of collapsing it to a single misleading verdict.
See the difference
What an accidental AI-crawler block looks like in robots.txt, why it happens, and the one-line fix.
Who needs to check AI-crawler access
- Site owners migrating hosts or CDNs
- Confirm a new hosting panel, WAF, or CDN didn't ship a default robots.txt that blocks AI crawlers by name.
- A line-by-line answer for each of 10 named crawlers, instead of guessing from a hosting provider's changelog.
- Teams that fought a bot-traffic spike
- Check whether an emergency 'block all scrapers' rule swept up PerplexityBot, ClaudeBot or GPTBot along with the abusive traffic.
- Confirmation of exactly which legitimate AI crawlers got caught in a defensive rule that was written in a hurry.
- Anyone who launched from a staging environment
- Verify a blanket 'Disallow: /' used during development was actually removed before the site went live, not just forgotten.
- A clean, dated answer instead of trusting memory about whether the dev-only rule was reverted.
- Content and SEO teams tracking AI visibility
- Add crawler access as a standing item in a pre-publish or quarterly technical check, the same way they already check indexability.
- A fast, free, repeatable check that quotes the exact rule so a developer can act on it without re-deriving the robots.txt logic themselves.
How the crawler check works
- 01
Paste your URL in @vustSEObot
No signup, no email, no install — the bot fetches your site's robots.txt and the target page directly over Telegram.
- 02
Each of 10 named crawlers is checked individually
GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, Google-Extended, CCBot, Applebot-Extended and Amazonbot are matched against the parsed robots.txt rules one at a time — never collapsed into a single pass/fail score.
- 03
You get the exact rule quoted back
If a crawler is blocked, the report shows the literal robots.txt line responsible (for example `Disallow: /`) and which user-agent group it fell under, so you can verify it yourself in seconds.
- 04
Meta-robots is checked too
Beyond robots.txt, the bot also reads the page's meta-robots tag for a stray `noindex`/`nofollow` that would undercut the fix even after robots.txt is corrected.
Check your robots.txt right now
Paste your URL in @vustSEObot — it checks 10 named AI crawlers in one pass and quotes the exact blocking rule, if any. Free, 10 checks/day, no signup.
Honest about what this check does and doesn't tell you
Every finding is a quoted, verifiable fact
The report never says 'this looks blocked' — it quotes the exact robots.txt line and the user-agent group that matched, so you can open the file yourself and confirm it in ten seconds. That's the entire trust anchor of this check.
Blocked ≠ invisible to that AI product
GPTBot governs OpenAI model training and some retrieval; it is not the single switch for every way ChatGPT might reference a page, since live browsing/search fetches can use a different crawler some sites never think to block. This tool tells you what the rule says, not what every downstream AI product does with that signal.
Allowed ≠ guaranteed citation
Letting every crawler read your site is necessary, not sufficient — an allowed crawler still has to find the page, parse it, and judge it worth citing. This check answers the access question honestly and stops there; no tool can promise the citation on top of it.
Free, fast, and narrow on purpose
This is one Tier-1 mechanical check, not the whole GEO picture. For structured data, answer-block presence, chunkability and freshness in the same pass, the full audit at @vustSEObot covers all of it — this page exists because the crawler-access finding is common, silent, and worth checking on its own.
Frequently asked questions
How do I check if GPTBot is blocked on my site?
Fetch yourdomain.com/robots.txt and look for a User-agent: GPTBot block followed by Disallow: / (or a path that covers your content). If there's no GPTBot-specific group and no blanket User-agent: * disallow, GPTBot is allowed by default. @vustSEObot automates this exact check — paste your URL and it quotes back the specific rule, if any, that blocks it.
Does blocking GPTBot mean ChatGPT can't cite my site?
Not exactly, and this is the part most "just add Disallow" advice gets wrong. GPTBot primarily feeds OpenAI's model training and some retrieval use — but ChatGPT's live browsing and citation behavior mostly runs through two SEPARATE user-agents, OAI-SearchBot and ChatGPT-User, that a site can leave wide open even while blocking GPTBot. @vustSEObot checks all three (GPTBot, OAI-SearchBot, ChatGPT-User) individually, not just GPTBot, precisely because they control different things.
What does Google-Extended actually control?
Google-Extended is a separate control for using site content to train Gemini models and to ground responses in Gemini Apps and Vertex AI. It does not control inclusion or ranking in Google Search, AI Overviews, or AI Mode. Sites may block it deliberately while leaving normal Googlebot crawling unchanged.
Why would a site accidentally block CCBot or PerplexityBot?
Three common causes: a default robots.txt shipped by a CMS, hosting panel, or security plugin that blocks a long list of bots by name; an overzealous "block all scrapers" rule added during a bot-traffic spike that swept up legitimate AI crawlers too; or a blanket Disallow: / added during development that was never reverted before the site went live. None of these are usually a deliberate decision to keep Perplexity or the Common Crawl out.
Is there a free tool to check all the AI crawlers at once?
@vustSEObot on Telegram checks 10 named AI crawlers in one pass — GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, and Amazonbot — and reports each individually as allowed or blocked, quoting the exact robots.txt line responsible. Free, up to 3 checks a day, no signup or email required.
Does allowing every crawler guarantee my page gets cited by AI answers?
No. Crawler access is a precondition, not a guarantee — a crawler that's allowed to fetch your page still has to find it, parse it, and judge it worth citing. This check only answers the access question honestly; it does not promise a citation.
More of the GEO/AEO audit
Ready when you are
Find out if AI answer engines can even read your site.
GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, Amazonbot — checked individually, with the exact robots.txt line quoted back. Free, in Telegram, no signup.