// Answer

How do I let AI crawlers access my site?

To let AI crawlers access your site, allow their user agents in robots.txt, avoid blocking them at the CDN or firewall, and confirm key pages return 200. Here is the step-by-step.

The short answer

To let AI crawlers access your site, check that robots.txt does not disallow the AI user agents you want to reach you (for example GPTBot, ClaudeBot, PerplexityBot, and Google-Extended), then confirm nothing at your CDN, WAF, or firewall is silently blocking them. Allowing an agent means either leaving it unlisted under a broad allow or adding an explicit Allow rule for it. After that, verify your important pages return a 200 status to those agents and are not gated behind JavaScript, login walls, or aggressive bot challenges.

Step one: audit robots.txt

Open yoursite.com/robots.txt and read it as the crawlers do. Look for any User-agent block that names an AI agent with a Disallow rule, and for broad Disallow rules that would catch them. If you want an agent to read your site, make sure no rule blocks the paths you care about. A single blanket Disallow or a blocked agent block is the most common reason a site never shows up in AI answers.

Step two: check the edge and the firewall

robots.txt is a request to well-behaved crawlers, but your CDN, WAF, or bot-management layer can block an agent before it ever reaches robots.txt. Review those rules for blocks keyed on user agent or IP. Then confirm your key pages return a real 200 with HTML content, not an empty shell that only fills in through JavaScript, and that no login wall or hard bot challenge sits in front of them.

Step three: verify, do not assume

Access is the first requirement for being cited, because an engine cannot quote a page it cannot fetch. Rather than guessing, test it. The free AI visibility scan checks AI-crawler access across 16 agents, reads your robots.txt and llms.txt, and reports exactly which engines can and cannot reach you.

See the AI citation tracking methodology for how access turns into citations, or what is AEO for the full practice.

Frequently asked questions

Which AI user agents should I allow?

The commonly discussed ones include GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot (Perplexity), and Google-Extended (which governs Google AI training and some AI surfaces). If you want to appear in AI answers, these should be able to reach your pages. You can allow them broadly in robots.txt or add explicit Allow rules, and you should confirm your CDN or firewall is not blocking them independently of robots.txt.

My robots.txt looks fine but crawlers still cannot reach pages. Why?

robots.txt is only one gate. A CDN, WAF, or bot-management rule can block an agent by user agent or IP before the request ever reaches your origin, and pages that render only through JavaScript or sit behind a login can look empty to a crawler. Check your edge or firewall rules, confirm key pages return a 200 with real HTML content, and avoid aggressive bot challenges on pages you want cited.

Run your free scan in 60 seconds

Run a free scan to see where you stand across ChatGPT, Claude, Gemini, and Perplexity: which answers cite you, which cite competitors instead, and what to fix first.

Run a free scan