// Methodology

AI Visibility Tracking Methodology

How AEO Goal measures AI visibility: prompt-set construction, engine coverage, crawl cadence, deterministic detection, the Share of Model formula, Wilson confidence intervals, exclusion rules, and limitations.

Quick Answer

AEO Goal measures AI visibility by running a stable, customer-defined prompt set across configured AI answer engines on a per-prompt cadence, detecting brand and competitor mentions deterministically against configured names and aliases, excluding placeholder and synthetic rows, and computing Share of Model as brand mentions divided by total tracked answers. Reports add a quality-weighted variant, a 95% Wilson confidence interval, and a prior-period delta so teams do not overreact to single-answer swings.

This page is written so an analyst can trace any headline number back to stored rows and recompute it. It covers the mention layer of measurement; the source-URL layer, including how cited URLs are verified before they count, is in the AI citation tracking methodology.

How Share of Model is measured: a stable tracked prompt set with market context runs on a per-prompt cadence across configured engines, deterministic detection records mentions, positions, and sentiment, exclusion rules remove placeholder and verified-false rows, and reporting computes Share of Model with a Wilson confidence interval and prior-period delta

How the prompt set is constructed

Every visibility metric is conditional on the prompt set, so the prompt set is explicit and customer-controlled rather than an opaque sample:

  • Prompts are buyer questions, defined per brand. A tracked prompt is a question a buyer would ask an answer engine, attached to one brand together with that brand’s names, aliases, and named competitors. Prompt research tooling can propose candidates from the brand’s category, competitors, and content, but nothing is tracked until a user adds it.
  • Prompts carry market context. Country, language, and locality settings can be applied per prompt, and market fields are persisted with every result row so visibility can be segmented by locale. Market allowances vary by plan and are listed on the pricing page.
  • Stability is the analyst’s control. The same prompt text reruns each cycle. Editing prompts mid-window changes the sample, and that, not engine behavior, is the most common cause of a surprising jump. Compare periods only across a stable prompt set.

Platform coverage

AEO Goal includes crawler adapters for ChatGPT, Claude, Perplexity, Gemini, Google AI Overview, Microsoft Copilot, Grok, Meta AI, and DeepSeek. These are code-supported platform paths; live coverage for a given account is always the subset with configured credentials and enabled crawler paths, and public reporting should keep that distinction.

When a provider key is missing or a call fails, the adapter returns a stub row that is skipped before persistence. Synthetic fallback rows, such as synthetic Google AI Overview fallbacks, are excluded from customer-facing Share of Model and citation-performance math. Failure therefore reduces sample size instead of injecting empty answers, which is the honest direction for the error to point.

Crawl cadence

Visibility is measured on a schedule, not in real time. Each tracked prompt has its own configured check frequency, with supported cadences from hourly through monthly, and paid recurring workflows run on separate scheduling paths from trial and onboarding flows. Failures are isolated per prompt: a failed prompt retries on a later cycle so one provider outage does not block the rest of a customer’s queue.

For auditing, the implication is that per-engine row counts inside a window can legitimately differ. A provider outage, a mid-window prompt addition, or a cadence change all alter the denominator, and each is visible in the stored rows.

Brand and competitor detection

Detection is deterministic and auditable. Each captured answer is matched against the brand names and aliases and the competitor names configured on the tracked setup. The detector records whether the brand appeared, the mention count, which competitors appeared, sentiment, and position: list rank when the answer is a structured list, sentence position when it is prose. A genuine literal match is required; an answer that only echoes the prompt or explicitly states the brand is absent does not count as a favorable mention.

This is deliberately not full NLP entity resolution or trademark disambiguation. A misspelling or an unconfigured alias will be missed, and completing the alias list in the tracked setup is the customer-side half of detection quality. The trade is intentional: a rule a team can inspect and reproduce beats a model verdict it cannot.

The Share of Model formula

Share of Model is the percentage of tracked AI answers that mention the brand:

Metric Calculation
Brand mentions Count of tracked answers where the brand was detected.
Total queries Count of tracked answers in the selected window after exclusions.
Share of Model (brand mentions / total queries) * 100.
Weighted Share of Model Quality-weighted average where non-mentions contribute zero and mentioned answers are weighted by citation position and sentiment.

As a worked example with illustrative numbers: if a window contains 200 tracked answers after exclusions and the brand was detected in 50 of them, Share of Model is 25%. If most of those mentions were late in the answer or neutral in tone, the weighted variant lands below 25%, because weighting rewards early, positive placement and gives non-mentions zero.

The exclusions in the denominator are the ones defined above and in the citation methodology: stub rows never persist, synthetic fallbacks are excluded, and citations explicitly verified false are removed while pending verifications remain counted until they resolve.

Confidence intervals and movement

Reports can include a 95% Wilson confidence interval around Share of Model and a delta against the immediately prior equal-length window. The Wilson interval is used because it behaves sensibly at small sample sizes and near 0% or 100%, which is exactly where prompt-level visibility data lives.

The reading rule is simple: movement inside the interval is indistinguishable from sampling noise; movement beyond it, sustained across cycles, is a trend. Small prompt sets produce wide intervals, which is the honest way to say that ten prompts cannot support a percentage-point-level conclusion. Teams that need tighter intervals should add prompts or lengthen windows rather than reading more precision into the number than the sample contains.

Market and locale handling

Tracked prompts run with persisted market context, so visibility can be segmented by locale at national, state, or city granularity with country and language settings. The crawl task applies the market context to the prompt and writes the market fields onto the resulting rows; segmentation at read time is a filter over those fields, not a re-estimate. Plan limits govern how many markets an account can configure.

Data included in reporting

Each tracked answer row carries the fields needed to recompute any aggregate:

Field Purpose
AI platform The answer engine or crawler path that produced the answer.
Model version The model or crawler version label recorded with the response.
Brand mentioned Whether the tracked brand appeared, and the mention count.
Competitors mentioned Which configured competitors appeared in the answer.
Citation position List rank or sentence position where the brand appeared.
Sentiment and quality score Signals used to weight answer quality.
Market fields Locale context persisted for market-level segmentation.
Source URLs URLs cited or extracted from the answer, when available.

Because rows are tenant-scoped and self-describing, an analyst with an export can rebuild Share of Model, the weighted variant, and competitor overlap for any window and filter, which is the practical test of a reproducible method.

Method versioning

The lastUpdated date on this page is the method version marker. When a formula, detection rule, or exclusion rule changes in a way that could move reported numbers, this page is updated to describe the new behavior. Aggregates are computed from stored rows at read time, so a rule change alters how history is aggregated, not the captured rows themselves; analysts comparing long ranges should read the current rules against their window.

Limitations

Stated plainly, because a visibility methodology that admits variance is more useful than one that promises certainty:

  • Non-determinism. The same prompt on the same engine can return different answers minutes apart. Single answers are anecdotes; the method is built for trends over stable prompt sets.
  • Sampling limits. Metrics describe the tracked prompts, not everything buyers ask. Wide Wilson intervals on small sets are the system telling you the sample is thin.
  • No control over engines. AEO Goal does not guarantee rankings, citations, or answer inclusion. Third-party AI platforms control their own responses and change them without notice.
  • Coverage is configuration. Credentials, enabled crawler paths, markets, and cadence all shape the sample. Period comparisons are only valid across stable configuration.
  • Alias-bounded detection. Deterministic matching misses names that were never configured.
  • Google AI Overview care. Real captured rows count; synthetic local fallback rows are excluded from customer-facing math and should never be read as measurement.

Enterprise recommendation

Treat AI visibility as a trend measurement, not a single-answer verdict. Use stable prompt sets, documented markets, repeatable cadence, configured-platform disclosure, verified citations, and confidence intervals before making strategy or procurement decisions. Read the AI citation tracking methodology for how cited sources are verified, the AI visibility tracking product page for how these metrics surface in the product, and AEO vs traditional SEO to frame Share of Model for stakeholders used to keyword rankings.

Frequently asked questions

Does AEO Goal claim every platform is always tracked?

No. AEO Goal has adapters for ChatGPT, Claude, Perplexity, Gemini, Google AI Overview, Microsoft Copilot, Grok, Meta AI, and DeepSeek, but live coverage depends on configured provider credentials and enabled crawler paths. Missing or placeholder provider responses are excluded from persisted customer-facing measurements.

Is AI visibility measured in real time?

No. AI visibility is measured on a scheduled cadence per tracked prompt, not in real time. Each prompt has its own configured check frequency, from hourly through monthly, and failed prompts retry on a later cycle without blocking the rest of the queue.

Why does AEO Goal report a confidence interval on Share of Model?

Because AI answers are non-deterministic and prompt sets are finite samples. The 95% Wilson interval expresses how much a Share of Model number could move from sampling alone, so a change inside the interval should be treated as noise, not a trend.

Run your free scan in 60 seconds

Run a free scan to see where you stand across ChatGPT, Claude, Gemini, and Perplexity: which answers cite you, which cite competitors instead, and what to fix first.

Run a free scan