Glossary

One defined term per page. Definitions are written to be quoted alone — that is the point.

AI CitationAn AI citation is a source link inside a generated answer. Pew data: ~1% of users click them — yet cited sites earned ~35% higher CTR (Seer, 2025).AI HallucinationAn AI hallucination is confident fabricated output, like the $249–$399 price ladder our pipeline extracted from a page carrying one $479 offer [our data].AI Share of VoiceAI share of voice is the share of tracked prompts where an answer engine mentions or cites your brand — 6 appearances on 25 prompts is 24% for the month.AI Training DataAI training data is the corpus a model learns from before it answers — not the index it searches at answer time. 2 corpora, 2 crawlers, 2 opt-outs.Answer EngineAn answer engine returns 1 composed, cited answer instead of ranked links. The named engines — and Pew's 2025 data on what that does to clicks.BingbotBingbot is Microsoft's crawler for the Bing index behind Bing search, Copilot and grounding API results. IndexNow pushes up to 10,000 URLs per request.ChunkingChunking splits pages into the passages AI engines index and quote. Why 2–4 sentence paragraphs and standalone sections are retrieval decisions, not style.Citable PassageA citable passage survives being quoted with zero context. The property, the 5 dependencies that break it, and the 2 blocks our contract enforces per page.Click-Through Rate (CTR)CTR is clicks divided by impressions. Ahrefs measured −34.5% CTR under AI Overviews (March 2025); Seer measured about +35% for cited pages (2025).Crawl-to-Refer RatioThe crawl-to-refer ratio counts crawler fetches per referred visit. Cloudflare's July 2025 data: ~38,000:1 for Anthropic, 5.4:1 for Google.EmbeddingAn embedding represents text as numbers so retrieval matches meaning, not words — why 1 page can answer questions it never uses the wording of.EntityAn entity is a thing — person, organization, product, concept — engines identify across sources. Why 1 consistent identity beats 100 keyword variants.GooglebotGooglebot is Google's crawler for the Search index that AI Overviews and AI Mode draw from. Google documents a 15MB default fetch limit and 3 ID signals.GroundingGrounding bases an AI answer on pages retrieved at answer time instead of training data alone; 24% of ChatGPT answers skip retrieval (Growth Memo, 2026).Information GainInformation gain is what a page adds that existing sources don't say. The Princeton GEO study's ~40% visibility lift is the closest measured cousin.JSON-LDJSON-LD is a block of JSON that describes a page to machines. It is 1 of the 3 structured-data formats Google documents, and the 1 it recommends.Knowledge GraphA knowledge graph stores entities and the typed relationships between them. What markup does and doesn't do, and why 1 stable identity beats many.Passage RetrievalPassage retrieval scores individual passages, not whole pages. Why the Princeton 2024 gains were passage-level, and what makes a passage standalone.Query Fan-OutQuery fan-out is how Google's AI features turn 1 query into many concurrent searches. Google documents it for both AI Overviews and AI Mode.SERP (Search Engine Results Page)A SERP is the page an engine returns for a query. Google now puts generated answers on it too, which is why 2 metrics stopped moving together.SnippetA snippet is the excerpt an engine shows or reuses from your page. Google documents 4 directives that limit it — across 6 result surfaces, AI included.Structured DataStructured data is machine-readable markup describing a page's meaning. Google documents 2 jobs for it, and a 1,885-page test found no AI-citation lift.Topical AuthorityTopical authority is measurable coverage depth on one subject — not a score any engine publishes. How a 167-page, five-cluster library demonstrates it.User-Agent StringA user-agent string is a client's self-reported ID. RFC 9309 matches only the product token inside it, case-insensitively, against 1 group of rules.User-Triggered FetcherA user-triggered fetcher requests 1 URL because a person just asked. Google names it 1 of 3 client categories; 3 AI vendors document 1 such agent each.Vector SearchVector search retrieves passages by meaning instead of matching words. Why 1 page can be retrieved for questions it never uses the words of.Web CrawlerA web crawler is an automated client that fetches URLs and follows links. RFC 9309 governs the rules; Google sorts its own clients into 3 categories.X-Robots-TagThe X-Robots-Tag carries robots meta tag rules in the HTTP response instead of the HTML. It is invisible in page source, and 1 stray noindex deindexes.