The short answer
Whether you should block AI crawlers depends on your goal, because the same block that protects content from training also keeps you out of AI answers. If visibility in ChatGPT, Perplexity, Gemini, and AI Overviews is a priority, you generally want the answer-facing agents allowed, since an engine cannot cite a page it cannot read. If protecting proprietary or paywalled content is the priority, blocking is reasonable. The practical middle path is selective: block the agents and paths you truly want withheld, allow the ones tied to answers, and decide per agent.
Understand the trade-off
Blocking an AI crawler is a real choice with a real cost. On one side, withholding your content from training and reuse can be the right call for proprietary research, paywalled articles, or material you simply do not want ingested. On the other side, the agents that assemble answers need to fetch your pages to cite them, so a broad block tends to remove you from the answers your customers now see first. There is no universally correct setting, only the one that fits your goal.
Block selectively, not by default
robots.txt supports per-agent and per-path rules, so you rarely need an all-or-nothing block. A common pattern is to allow the answer-facing agents (for example GPTBot, ClaudeBot, PerplexityBot, Google-Extended) on your public, marketing, and reference pages while disallowing sensitive paths like member areas or paywalled content for everyone. Confirm your CDN or firewall enforces the same intent, since an edge rule can block an agent even when robots.txt allows it.
Decide with real data
Before you block broadly, see what you would give up. The free AI visibility scan shows which agents currently reach your site and how it is exposed, and AI citation tracking shows the citations you would lose by closing a door.
See AEO vs traditional SEO for how this fits the wider shift from ranking to being cited.