Quick Answer
AEO Goal tracks AI citations by running a stable, customer-defined prompt set against configured AI answer engines on a per-prompt cadence, recording every answer with its source URLs, detecting brand and competitor mentions deterministically, filtering placeholder and synthetic provider rows, and verifying cited URLs in two stages before they influence customer-facing metrics. Citations verified false are excluded from Share of Model math.
This page exists so a skeptical analyst can audit the numbers. It describes what the product actually does, in the order the pipeline does it, including the parts that limit precision. The companion page, the AI visibility tracking methodology, covers the mention-level metrics; this page covers the citation and source-URL layer.
What counts as citation data
An AI citation is a recorded instance of an answer engine surfacing a source URL, brand mention, or competitor mention in response to a tracked prompt. One tracked answer produces one row, and each row stores tenant-scoped fields: the prompt text, the engine and model label that produced the answer, the market context the prompt ran under, the full answer text, whether the brand was mentioned and how many times, sentiment, position, quality signals, and any source URLs the engine exposed or that could be extracted from the answer.
Two distinctions matter for reading the data honestly:
- Mentioned is not cited. An engine can name the brand while building its answer from a third-party page, and it can use a brand’s page as evidence without naming the brand. The rows record both signals separately so the AI citation tracking product can report them side by side.
- Exposed is not universal. Some engines, such as Perplexity, expose citation links directly. Others provide answer text where source influence has to be read through URLs and mention patterns. The engine label on each row tells you which kind of evidence you are looking at.
Rows are measurement data, not a guarantee that a third-party platform will keep citing the same source. Engines revise their answers constantly.
How the prompt set is constructed
Citation metrics are only as meaningful as the prompt set behind them, so the prompt set is customer-controlled and versioned by ordinary edit history rather than sampled opaquely:
- Prompts are defined per brand. Each tracked prompt is a buyer-style question attached to a specific brand, with the brand’s names and aliases and its named competitors attached to the same tracked setup. Prompt research tooling can suggest candidate prompts from the brand’s category and content, but suggestions become tracked prompts only when a user adds them.
- Prompts carry market context. A prompt can run with country, language, and locality context, and those market fields are persisted on every resulting row so citation reporting can be segmented by locale. Plan limits control how many markets an account can configure.
- Prompts are stable between edits. The system reruns the same prompt text on each cycle. If a team edits a prompt, subsequent rows reflect the new text; the rows captured under the old text remain in history. Analysts comparing periods should confirm the prompt set did not change mid-window, because a changed prompt set is the most common non-engine cause of a metric jump.
Per-engine execution
AEO Goal includes crawler adapters for ChatGPT, Claude, Perplexity, Gemini, Google AI Overview, Microsoft Copilot, Grok, Meta AI, and DeepSeek. Implemented coverage is wider than live coverage: a platform only produces rows when its provider credentials are configured and its crawler path is enabled. When a key is missing or a provider fails, the adapter returns a stub that is skipped before persistence, so unavailable providers produce no rows rather than empty ones.
Execution is scheduled, not real time. Each prompt has its own configured check frequency, with supported cadences from hourly through monthly. On each cycle, every due prompt runs against every configured engine, and failures are isolated at the prompt level: a failed prompt is retried on a later cycle without blocking the rest of the queue. This is important for auditing because it means row counts per engine can legitimately differ within a window when one provider had an outage or a prompt was added mid-window.
How mentions and citations are detected
Detection is deterministic and auditable by design. Each answer is matched against the brand names and aliases and the competitor names configured on the tracked setup. The detector records whether the brand appeared, how many mentions were found, which competitors appeared, and where the brand appeared: list position when the answer is structured as a list, sentence position when it is prose. Detection requires a genuine literal match in the answer text; an answer that merely echoes the prompt or states that the brand is absent does not count as a favorable mention.
This is intentionally not full NLP entity resolution. It will not catch a misspelled brand or an alias nobody configured, and it says so here rather than pretending otherwise, because a deterministic rule a team can inspect beats a fuzzy one it cannot. Supplying complete aliases in the tracked setup is the customer-side half of detection quality.
Placeholder and synthetic response filtering
Two classes of rows are kept out of customer-facing math:
- Stub rows from missing credentials or failed provider calls, skipped before persistence.
- Synthetic fallback rows, such as synthetic Google AI Overview fallbacks used in non-production paths, excluded from customer-facing Share of Model and citation-performance aggregation.
The purpose is honesty under failure: an unavailable provider should reduce sample size, not fabricate zeros or fill the window with empty answers.
Two-stage citation verification
Answer engines sometimes cite URLs that do not resolve or pages that never mention the brand. Verification confirms a cited URL is real evidence before it influences customer-facing metrics:
| Stage | What happens | Failure outcome |
|---|---|---|
| Reachability | A timed HEAD request checks the cited URL. | Non-success or unreachable URLs are marked unverified, with the error recorded. |
| Body and brand check | The page body is fetched, visible text is extracted, and a strict judge checks whether the page genuinely mentions the brand. | Rows that fail are marked verified false and excluded from Share of Model math. |
The verifier records its method, the HTTP status, whether the page mentions the brand, a confidence signal, a supporting snippet when available, and an error field when verification fails, so every verdict can be inspected later.
Share of Model inclusion rules
Customer-facing Share of Model counts citations in the selected date window only when the row is not a synthetic fallback and not explicitly verified false. Rows with verification still pending remain in the calculation until the check completes. That decision rule is deliberate: dropping fresh citations to zero while the verifier catches up would understate real visibility, while counting confirmed-false citations would overstate it. The pending window is the accepted trade-off, and it is why a number can tick down slightly after a scan as late verifications resolve.
Quality and position signals
Citation quality is a weighted signal built from mention status, sentiment, and position. A brand mentioned early and positively contributes more than a late or neutral mention, and non-mentions contribute zero to weighted Share of Model. Position is recorded transparently as list rank or sentence position so a team can audit exactly why a weighted number moved. These signals feed the AI citation tracking dashboard and competitor AI visibility comparisons.
Data retention
Account deletion uses a 30-day soft-delete grace period before hard deletion cascades through associated user data. This methodology page makes no claim that prompt or answer rows auto-expire on a shorter TTL; retention beyond deletion behavior is a contractual matter, not a marketing one. Security and data-handling practices are documented on the security page.
Method versioning
The lastUpdated date on this page is the method version marker. When an inclusion rule, detection rule, or verification stage changes in a way that could move customer-facing numbers, this page is updated to describe the new behavior, and teams comparing long time ranges should read the current rules against the window they are analyzing. The formulas described here are computed from stored rows at read time, so rule changes apply to how rows are aggregated, not by rewriting history.
Limitations
A methodology page that hides its error sources is not a methodology page. Known limits:
- Answer engines are non-deterministic. The same prompt on the same engine can produce different answers minutes apart. Single rows are anecdotes; trends over a stable prompt set are the signal.
- Sampling is bounded by the prompt set. Metrics describe the tracked prompts, not the full space of questions buyers ask. A prompt set that misses a key buyer question misses its citations too.
- Verification depends on the third-party web. A page can change or disappear after the answer was captured, HTTP behavior varies, and visible-text extraction is imperfect on heavily scripted pages.
- Alias coverage is customer-supplied. Deterministic detection misses names that were never configured.
- Coverage depends on configuration. Engine credentials, enabled crawler paths, markets, and cadence all shape the sample. Compare periods only when configuration was stable.
- A citation is not the whole story. A source can shape category context without naming the brand, so teams should read verified source URLs and broader answer framing together.
Enterprise recommendation
Treat citation data the way an analyst would treat any sampled measurement: use repeated runs over a stable prompt set, verified citations only, documented markets, and configured-platform disclosure before drawing conclusions. Do not present a single AI answer or one unverified URL as procurement proof. Pair this layer with competitor AI visibility benchmarking to see how verified citations translate into Share of Model against rivals, and read how AEO differs from traditional SEO when framing these metrics for stakeholders used to rankings.