The short answer
AI engines usually skip a site for one of three reasons: they cannot reach it, they cannot cleanly extract an answer from it, or they do not treat it as authoritative on the topic. Access problems come from robots.txt blocks, CDN or firewall rules, login walls, or content that renders only through JavaScript. Extraction problems come from buried answers, vague prose, missing headings, and absent JSON-LD. Authority problems come from thin content and weak entity signals on the subject.
Rule out access first
An engine cannot cite a page it cannot fetch, so this is always the first check. Look for AI user agents disallowed in robots.txt, blocks at your CDN or WAF keyed on user agent or IP, pages behind a login, and content that only appears after JavaScript runs. Confirm your key pages return a 200 with real HTML to the crawlers. If access fails, nothing else you do will help.
Then check extractability
If crawlers can reach you but still do not cite you, the answer may be too hard to lift. Engines favor pages that state the answer directly near the top, use clear headings, and label facts with JSON-LD. Long preambles, hedged prose, and answers scattered across a page all raise the cost of extraction and push an engine toward a cleaner source.
Then look at authority
If access and extraction are fine, the remaining gap is usually trust. Engines lean on sources they treat as credible on a topic, which comes from real depth, corroborating pages, and clear entity signals, not a single thin page with the right keywords.
Diagnose it rather than guess. The free AI visibility scan checks AI-crawler access, llms.txt, JSON-LD, metadata, and entity salience, and AI citation tracking shows how often engines actually mention you versus competitors.