Crawlability

Crawlability helps you understand which AI crawlers are allowed or blocked from accessing your website based on your robots.txt configuration.

Simply select one of your tracked domains, and Cite AI automatically analyzes your robots.txt file against supported AI crawlers. No additional setup is required.

This helps you identify whether important AI platforms can crawl your content and whether any robots.txt rules are unintentionally limiting your AI visibility.

Crawlability Overview

Cite AI evaluates your robots.txt rules for every supported AI crawler and displays their crawl permissions.

For each crawler you’ll see:

FieldDescription
BotThe crawler’s user-agent (for example, GPTBot or ClaudeBot).
PlatformThe AI provider behind the crawler, such as OpenAI, Google, Anthropic, or Perplexity.
Bot TypeThe crawler’s primary purpose, including Training, Search, User Retrieval, or Other.
StatusWhether the crawler is Allowed, Partially Allowed, or Blocked.
ReasonWhich robots.txt rule determined the result, including explicit bot rules or inherited wildcard rules.

Use the available filters to search by platform, bot type, or crawl status.

URL Tester

The URL Tester allows you to verify crawl permissions for any page on your website.

Simply enter a URL from your domain to instantly see which AI crawlers can access it and which are blocked.

This is useful for:

  • Verifying robots.txt changes.
  • Troubleshooting crawl issues.
  • Confirming important pages remain accessible to AI search engines.

Understanding Crawlability

If an AI crawler is blocked by your robots.txt file, it cannot access the content on that page.

Blocked content may be less likely to:

  • Be crawled for AI search.
  • Be retrieved during AI browsing.
  • Be referenced in AI-generated responses.

Use Crawlability to:

  • Identify accidental crawler blocks.
  • Verify robots.txt changes before deployment.
  • Understand which AI platforms can access your content.
  • Improve website accessibility for supported AI crawlers.

Bot Categories

Cite AI groups supported crawlers based on their primary purpose.

Training Bots

Training bots collect publicly available content that may be used to train or improve AI models.

Examples include:

  • GPTBot
  • ClaudeBot
  • Google-Extended
  • DeepSeekBot
  • Meta-ExternalAgent
  • Cohere AI
  • Applebot-Extended

Search Bots

Search bots retrieve live web content to answer user queries and generate up-to-date responses.

Examples include:

  • OAI-SearchBot
  • Claude-SearchBot
  • PerplexityBot
  • AzureAI-SearchBot
  • Google Cloud Vertex Bot
  • Applebot
  • Amazon SearchBot

User Retrieval Bots

These bots visit websites while responding to a specific user’s request inside an AI assistant.

Examples include:

  • ChatGPT-User
  • Claude-User
  • Gemini Deep Research
  • Perplexity-User
  • Google-Agent
  • MistralAI-User
  • NovaAct

Other Bots

Some supported crawlers perform specialized functions that don’t fall into the categories above, such as research, indexing, metadata collection, or platform-specific services.

These bots are grouped under Other and are monitored separately within Crawlability.

By understanding which AI crawlers can access your website, you can ensure your content remains discoverable across the AI platforms that matter most while maintaining full control over your robots.txt configuration.