Skip to content

Reference

The vocabulary of AI search, defined without hedging

35 terms covering measurement, crawler access, sources, process and commercial models. Written so a definition can be quoted on its own and still make sense — which is the same standard the product holds your content to.

  • 35 defined terms
  • 6 categories
  • Published as DefinedTermSet markup
  • Quotable definitions

Start here

The three questions behind most searches

What is GEO in one sentence?

Generative engine optimization is the practice of being the source a generative AI engine retrieves, reasons over and cites when it answers a buyer's question — as opposed to competing for a position in a list of links.

What is AEO in one sentence?

Answer engine optimization is structuring content so a single passage answers a single question completely and can be extracted as the answer, with the markup and corroboration an engine needs to trust it.

How is AI share of voice different from SEO rank?

SEO rank describes a position in a ranked list a human scans. AI share of voice describes how often and how favourably a brand appears inside synthesised answers that cite only a handful of sources. There is no page two, so absence is total.

01 · 4 terms

Core concepts

The vocabulary of buying a recommendation instead of a ranking. These four terms are the ones most often used loosely, which is exactly why they are worth defining first.

Generative Engine Optimization (GEO)

GEO
GEO is the practice of making a brand more likely to be retrieved, reasoned over and cited by a generative AI engine. Where SEO competes for a position in a ranked list, GEO competes to be part of the source material an engine synthesises its answer from.
GEO levers that matter most: crawler access, entity clarity, machine-readable structure, transparent facts, and third-party corroboration on the domains the engine already trusts.

Answer Engine Optimization (AEO)

AEO
AEO is the practice of formatting content so it can be extracted as a direct answer — a self-contained passage that answers one question completely, with the supporting structure an engine needs to trust it.
In practice: question-shaped headings, a forty-to-sixty word answer immediately beneath, FAQPage or HowTo markup, and no dependency on surrounding context to make sense.

AI search visibility

AI visibility
AI search visibility is the degree to which conversational AI surfaces name and recommend your brand when buyers ask category questions. It is measured across two dimensions: whether you appeared at all, and how favourably you were framed.
The failure mode is silent. An omitted brand receives no impression, no click, and no analytics event — the buyer simply never learns the brand exists.

The shortlist problem

Shortlist
A conversational engine answers a buying question with two or three named options plus a rationale. There is no page two. A brand outside that shortlist is excluded from the evaluation entirely, which is why share of voice matters more here than average rank.

02 · 8 terms

Measurement

How visibility becomes a number. Every definition here maps to a formula the platform publishes rather than to an internal opinion.

AI share of voice

Share of voice
AI share of voice is the proportion of tracked buyer prompts on which a brand is cited by an AI surface. It is an unweighted presence measure: citing answers divided by total checks, multiplied by one hundred.
Read it alongside the weighted visibility score. A brand can hold high share of voice while being consistently framed as a fallback, and that gap is itself the finding.

Weighted visibility score

Visibility score
A position-weighted and stance-weighted measure of how well a brand was positioned across every tracked check, normalised to zero through one hundred so it is comparable across brands and across weeks.
Position weight is one divided by log base two of position plus one. That is multiplied by a stance multiplier between zero and 1.25, summed, then normalised against the maximum achievable in a complete cycle.

Position weight

Position
An ordinal discount applied by where a brand appeared in an answer. A first mention scores one, a second mention roughly 0.63, a third roughly 0.50. Logarithmic decay prevents a long tail of low mentions from outweighing one genuine recommendation.

Citation stance

Stance
How an engine framed a brand within an answer: first choice, recommended, alternative, mentioned only, or cautioned against. Stance is separate from sentiment — an engine can describe a product positively while positioning it as the fourth-best option.

Stance multiplier

Multiplier
The weight applied on top of position weight for each stance: 1.25 for first choice, 1.0 for recommended, 0.75 for alternative, 0.5 for mentioned only, and zero for cautioned against.
Cautioning scores zero rather than negative. It cannot contribute at any position, and separately raises a critical incident — because a caution needs a response, not an accounting adjustment.

Sentiment

Sentiment
The emotional valence of what an engine said about a brand, classified as positive, neutral or negative. Negative sentiment is the trigger for the hallucination signal, because a negative description is frequently factually wrong rather than merely unflattering.

Citation drift

Drift
The change in how AI engines describe a brand between measurement cycles. Drift happens without any action by the brand, because models are re-indexed continuously and the third-party sources beneath an answer change independently.
Drift detection compares each answer to the previous processed answer for the same prompt and surface, which is why a monthly snapshot hides the incidents a weekly one catches.

Omission rate

Omission
One hundred minus the decision-stage visibility score: the proportion of buying-stage checks where the brand was absent. It is the input to the missed-shortlist forecast.

03 · 8 terms

Access and structure

The technical surface. If an engine cannot reach a page or cannot parse what it means, no amount of content quality matters.

Retrieval crawler

Retrieval bot
A crawler that fetches pages to answer a live or indexed query, as opposed to one that collects training data. Retrieval crawlers are the ones that affect what an engine says about a brand this week rather than what a future model knows.
OpenAI separates GPTBot (training) from OAI-SearchBot (indexing) and ChatGPT-User (user-initiated). Anthropic separates ClaudeBot, Claude-SearchBot and Claude-User. Blocking the training bot does not remove the brand from live answers, and blocking the search bot silently does.

Google-Extended

Google-Extended
A robots.txt token that opts a site out of Gemini training use. It is not a crawler and does not affect Google Search, AI Overviews or AI Mode, which all reach content through the regular Googlebot.
This is the most widely misunderstood directive in the category. Disallowing Google-Extended does not remove a brand from Google AI surfaces; only leaving Google entirely would do that.

llms.txt

llms.txt
A proposed convention that places a curated Markdown index at a domain root, giving language models an explicit map of the pages that matter instead of forcing them to infer hierarchy from navigation markup.
It is not a ratified standard and no major provider has publicly committed to it. It costs almost nothing to publish and some agents already read it, so the expected value favours adopting it — as a low-cost signal, not a guaranteed ranking factor.

llms-full.txt

llms-full.txt
An optional companion to llms.txt that carries the complete corpus in Markdown rather than a curated index, for agents that want full depth rather than navigation.

AI crawler access audit

Crawler audit
A check of robots.txt directives per AI user agent, using longest-prefix matching with the wildcard group as fallback, reporting each bot as explicitly allowed, explicitly blocked, or unspecified.
Unspecified is not the same as blocked. A site with no robots.txt at all defaults to full crawler access, which is why a blocking directive is the only finding that is genuinely urgent.

Entity disambiguation

Entity clarity
Making it unambiguous which real-world organisation a brand name refers to, by declaring a canonical Organization entity and linking it to authoritative profiles through sameAs. This is what stops a model merging a brand with an unrelated company of the same name.

Semantic structure

Semantics
The use of meaningful HTML — one H1, ordered headings, tables or definition lists, landmark regions and sufficient visible text — so a parser can determine what a page is about without guessing from styling.

JSON-LD structured data

JSON-LD
A script-tagged JSON format that describes page content in Schema.org vocabulary. It is the structured data format search and answer engines parse most reliably, because it is separate from presentation markup.

04 · 5 terms

Sources and authority

Answers are assembled, not retrieved from one page. These terms describe who actually decides what an engine says about a brand.

Citation source

Source
A domain an AI surface cites or draws on when answering a query. Sources are the real competitive battleground: an engine recommends what its trusted sources describe, so a brand absent from those sources is absent from the answer.

Source gap

Source gap
A domain that feeds answers in a category, carries competitors, and does not mention the brand. Source gaps are the highest-leverage remediation list, because they are specific, finite and directly attributable to the omission.

Third-party corroboration

Corroboration
Independent evidence about a brand on domains the engine already trusts — review hubs, community threads, editorial comparisons, documentation. Models weight corroboration over self-description, which is why a perfect product page cannot outrank a well-cited competitor profile.

AI hallucination about a brand

Hallucination
A confidently stated claim about a brand that is factually wrong — pricing that does not exist, a compliance certification never held, a feature removed years ago. It is detected through negative sentiment combined with a forensic gap reason.
The response is a factual correction kit: a direct refutation, FAQPage markup asserting the documented facts, a verified-facts block for llms.txt, and a re-run of the cycle to confirm the correction landed.

Corroboration-weighted content

Weighted content
Content written so that it is useful to the third party publishing it, not only flattering to the brand. Outreach that reads as promotion is discarded by both moderators and models; specific, comparative, non-promotional contributions are what get retained and cited.

05 · 6 terms

Process and proof

How the work is sequenced, and how its effect is established without dressing up correlation as causation.

Tracking cycle

Cycle
One ISO week of checks across every tracked prompt and every AI surface, stored as a single snapshot. Cycles are the unit of comparison: a delta always refers to two named cycles.

Grounded buyer prompt

Grounded prompt
A query written from crawled evidence about a brand and its competitors — real pricing structure, real feature deltas, real constraints — rather than from a keyword list, so it resembles what a buyer would actually type into an AI engine.

Forensic diagnosis

Forensics
The extracted explanation of why an outcome occurred: the winning competitor, the reason the engine gave, the brand-specific gap, and a verbatim quote from the answer. It converts a score change into an investigation with a starting point.

Baseline lock

Baseline
Freezing the current visibility score at the moment an action is marked complete, so its effect can be measured against the state that genuinely existed rather than against a moving reference.

Measured lift

Measured lift
The difference between a locked baseline and the value recorded on a later complete cycle. Until that later snapshot exists, the action reports as awaiting measurement rather than as a win.

Cycle completion rate

Completion
The share of a cycle’s checks that produced a scorable result. A partial cycle reads conservatively because per-surface scores are normalised against the full prompt count, never against the subset collected so far.

06 · 4 terms

Commercial

The terms used when the work is bought rather than built.

Software with a Service (SwaS)

SwaS
A model that sells a measurement platform and an implementation service on the same data, so the instrument and the crew that acts on it are not separated into two vendors blaming each other.

Managed GEO sprint

Sprint
A done-for-you engagement in which an engineer implements the full remediation queue — answers, llms.txt, entity schema, comparison pages and citation outreach — and reports the measured result against each locked baseline.

Remediation tier

Tier
One of five ranked classes of fix, ordered by time to value: a five-minute objection-buster answer, an llms.txt production asset, entity disambiguation, a two-hour versus blueprint, then ongoing citation authority outreach. Each fires only when its trigger is observed in your own data.

Model Context Protocol (MCP)

MCP
An open protocol that lets an AI client call tools exposed by a service. The platform exposes brand-scoped visibility tools over MCP, so an engineering agent can pull citation reports, prompts, sources and the action queue directly into its own workflow.

Where these terms are used

From definition to implementation

Each of these concepts maps to something concrete in the product rather than to a slide.

Share of voice, measured

Both numbers — weighted visibility and unweighted citation share — reported side by side with the formulas published.

Crawler access, audited

Per-bot robots.txt findings for GPTBot, ClaudeBot, PerplexityBot and Google-Extended, each with a copy-ready fix.

Drift, caught weekly

Five severity-ranked signals including the five-point macro drop and the decision-stage conquest.

Measured lift, not asserted

A locked baseline at completion and a measured after-value on the next complete cycle.

MCP, so agents can read it

Six brand-scoped tools over Streamable HTTP, authenticated with per-brand bearer tokens.

Definitions, questioned

The arguments worth having

Also published as FAQPage structured data, so an answer engine quotes the qualified version rather than an oversimplification.

Looking for the product rather than the vocabulary? See the capability surface.

Is GEO just a rebranded version of SEO?

No, though they share foundations. SEO competes for rank in a list a human scans. GEO competes to be the source a model retrieves, reasons over and cites inside a synthesised answer. Entity clarity, crawler permissions, machine-readable structure and third-party corroboration carry far more weight in GEO than link ranking does.

Does llms.txt actually affect AI visibility?

It is a proposed convention rather than a ratified standard, and no major provider has committed to it publicly. It costs almost nothing to publish and it makes your hierarchy explicit for any agent that reads it, so the expected value favours publishing it — but treat it as a low-cost signal, not a guaranteed ranking factor.

What is a citation stance and why weight it?

Stance describes how an engine framed you: first choice, recommended, alternative, mentioned only, or cautioned against. A mention is not a win — being listed fourth as a fallback is materially different from being named as the primary recommendation, so the score multiplies positional weight by stance weight.

What is citation drift?

Drift is the change in how engines describe you between cycles. Models are re-indexed and re-trained continuously, and competitor content and third-party sources shift underneath you. A brand can lose a first-choice position without any change of its own, which is why week-over-week comparison matters more than any single reading.

From vocabulary to evidence

Apply the definitions to your own citations

The concepts on this page are the ones the platform measures directly. Onboarding shows you where you actually stand against them in a single cycle.