Sign up free

AI Crawler Checker

See which AI crawlers your robots.txt lets in, and whether your site has an llms.txt.

About AI Crawler Checker

AI companies crawl the web to train models and to answer questions live. Use the AI crawler checker when you're deciding whether to block AI training, after a CDN or plugin has added AI rules for you, or when your site never shows up in AI search answers. Enter your site and click Check.

Your robots.txt is read for 25 AI crawlers, including GPTBot, OAI-SearchBot and ChatGPT-User from OpenAI, ClaudeBot and Claude-SearchBot from Anthropic, Google-Extended, Applebot-Extended, PerplexityBot, CCBot, Bytespider, meta-externalagent and Amazonbot. Each is labeled AI training, AI search or fetches for users, with allowed or blocked, the deciding rule and whether your file names it. The checker also looks for /llms.txt and /llms-full.txt and shows the start of llms.txt, and it spots noai or noimageai in meta robots or an X-Robots-Tag header. If AI search crawlers are blocked, you get a warning, because your pages could vanish from AI search answers. Write rules with the Robots.txt Generator and an llms.txt with the llms.txt Generator.

How to use AI Crawler Checker

  1. 1
    Enter your site

    Type the domain or any page address on the site.

  2. 2
    Click Check

    robots.txt, llms.txt, llms-full.txt and the home page are fetched.

  3. 3
    Read the table

    Each AI crawler shows its purpose, allowed or blocked, and the deciding rule.

  4. 4
    Decide and adjust

    Block training crawlers if you like, but keep AI search crawlers in if you want to be cited.

Why use Cubfile for this

  • 25 AI crawlers

    OpenAI, Anthropic, Google, Apple, Perplexity, Meta, Amazon, ByteDance and more.

  • Purpose labels

    AI training, AI search or fetches for users.

  • llms.txt check

    Finds llms.txt and llms-full.txt and shows how llms.txt begins.

  • noai signals

    Spots noai and noimageai in meta robots or X-Robots-Tag.

FAQ

AI Crawler Checker: questions and answers

What's the difference between AI training and AI search crawlers?
Training crawlers such as GPTBot and CCBot collect pages to build models. AI search crawlers such as OAI-SearchBot and PerplexityBot fetch pages to cite in answers. You can block one kind and allow the other.
Does blocking Google-Extended remove my site from Google Search?
No. Google-Extended only controls whether Google may use your content for its Gemini AI models. Crawling and ranking in Google Search are not affected.
What is llms.txt?
A proposed Markdown file at /llms.txt that gives AI assistants a short summary of your site and links to key pages. It's optional and not an official standard, but some AI tools read it.
Do AI crawlers obey robots.txt?
The major ones, including GPTBot, ClaudeBot and PerplexityBot, say they do. But robots.txt is a request, not a lock, so blocking at the server or firewall is the only hard stop. Before blocking by IP, verify crawler IPs so you don't shut out Google, Baidu or Bing.
Is the AI crawler checker free?
Yes, and it doesn't use your daily tasks. The files are fetched when you click, and nothing is stored.
Share AI Crawler Checker with a friendIt runs in any browser, and they can try it without signing up.

Related tools

SEO Broken Link CheckerCheck every link on a page and list the ones that are broken or redirected.
SEO SEO CheckerCheck a page on 40+ on-page and technical SEO points, with a fix for each problem.
SEO Indexability CheckerFind out whether a page can be indexed: status, robots, noindex and canonical in one check.
SEO Keyword Density CheckerSee which words and phrases a page repeats most, and how often.
SEO Meta Tag CheckerRead a page’s title, description, robots, canonical and social tags, with length checks.
SEO Robots.txt TesterRead a site’s robots.txt and test whether a URL is allowed for each crawler.
SEO Search Engine Spider SimulatorSee a page the way Baiduspider or Googlebot does, and whether it differs from what visitors get.
SEO Verify Googlebot and Baiduspider IPsCheck whether an IP in your logs really belongs to Google, Baidu or Bing.