AI crawler access is the most commonly overlooked AEO prerequisite. A robots.txt file that blocks GPTBot, ClaudeBot, or PerplexityBot produces exactly zero citations from that engine - regardless of content quality, FAQ schema, or off-site authority signals. This is a binary gate: pass it and citations become possible; fail it and citations are structurally impossible.

The AI Crawlers Your robots.txt Must Allow

These bots must be explicitly allowed (or not blocked) in your robots.txt for their respective engines to index and cite your content:

  • GPTBot - ChatGPT's web crawler (OpenAI)
  • OAI-SearchBot - ChatGPT search integration
  • ChatGPT-User - ChatGPT browsing
  • ClaudeBot - Claude's web crawler (Anthropic)
  • Claude-SearchBot - Claude search integration
  • PerplexityBot - Perplexity AI crawler
  • Google-Extended - Google AI training and Gemini
  • Googlebot - Standard Google crawling (required for Gemini's local data)

How to Audit Your Current robots.txt

Navigate to yourdomain.com/robots.txt. Look for any Disallow rules that match: a wildcard Disallow: / under any bot's User-agent (this blocks everything), rules blocking /blog/, /articles/, or other content directories where your AEO content lives, and User-agent: * blocks that do not explicitly re-allow the AI crawlers above.

A common problem: WordPress security plugins and CDN WAF rules that add blanket bot blocks. These silently block all AI crawlers and are often invisible in the robots.txt file itself.

The Correct Configuration

The safest configuration for maximum AI visibility allows all named bots explicitly, then uses Disallow rules only for pages you genuinely want excluded (admin panels, checkout pages, internal search results). Do not block AI bots to "protect" your content - blocking them means they cannot cite you, and unblocking them does not cause scraping or security issues.

Frequently Asked Questions

Should I block GPTBot to prevent OpenAI from training on my content? Blocking GPTBot for training data purposes also blocks ChatGPT from citing your content in browsed responses. If citations matter to your business, allow GPTBot. The training data concern is separate from citation visibility.

What happens if I accidentally block PerplexityBot? Your content produces zero citations from Perplexity. Users asking Perplexity about your service category in your city will never see your business named - regardless of how well you rank on Google.

How do I check if a specific bot is blocked? Use Google Search Console's URL Inspection tool for Googlebot. For other bots, use an online robots.txt tester with the specific bot User-agent string.