Resources / Glossary
AI SEO glossary
Concise definitions for AI SEO terms, including the vocabulary behind the site's six-step framework. Each entry links related terms and, where one exists, a resource that puts the concept into practice.
A
Agentic commerce Structured data
Shopping in which an AI assistant discovers products, compares them, and in some programs completes a purchase on the shopper's behalf. Because the assistant quotes a price and stock status rather than linking to a page the shopper reads, the merchant data behind the answer has to be current enough to transact on, not only accurate enough to describe.
Agentic Commerce Protocol (ACP) Structured data
An open standard, governed by OpenAI and Stripe as founding maintainers and licensed under Apache 2.0, that defines how an AI agent completes a purchase with a business. It covers the transaction through an agentic checkout API and a delegate payment API, and uses date-based version numbers. It is separate from OpenAI's product feed specification, which describes how the catalog reaches OpenAI in the first place.
AI answer
A direct response an AI system generates for a user's query, such as a chatbot reply or an AI Overview, instead of a ranked list of links. An AI answer can quote, summarize, or cite one or more sources, and a page can be used in an answer without producing a click.
AI citation Measurement
A displayed link or attribution that identifies a source associated with an AI answer. A citation does not by itself establish that the source supports every claim or that a reader clicked it. Track citations, brand mentions, and referral visits separately.
AI crawler Crawlability
An automated program used to collect web information for an AI service. Providers can use different agents for training, search indexing, and user-triggered retrieval. Where robots.txt applies, a wildcard group can cover unnamed agents; specific groups allow different policies. Verify each provider’s documented agent and control behavior.
AI Overview Measurement
An AI-generated answer feature in Google Search that can summarize information and display source links. Use the available Search Console report and its documented metric definitions to assess visibility. Website analytics does not necessarily identify whether an individual Google visit came from an AI Overview.
AI SEO
The practice of making a website's content and technical signals accessible, understandable, and citable to AI systems that crawl, retrieve, and generate answers, in addition to the search engines that rank pages. AI SEO extends established SEO practice; it does not replace the technical, structural, and editorial work ranking has always required.
Answer engine optimization (AEO)
Practices aimed at getting content selected, represented accurately, and cited inside AI-generated answers, as distinct from ranking in a list of links. The term is used close to interchangeably with generative engine optimization; this site treats AI SEO as the umbrella term for both.
Audit trail Auditing & correcting
A recorded sequence of what was observed, checked, decided, and changed while investigating an inaccurate, outdated, or manipulated AI answer. An audit trail lets a team show what happened if the same error recurs or a stakeholder asks how a correction was handled, rather than relying on memory of the case.
B
Brand mention Measurement
A reference to an organization, product, or person in an AI answer, with or without a link. Record mention presence separately from citations: an answer can name the brand, cite its website, do both, or do neither.
C
Canonical tag Crawlability
An HTML link element (rel="canonical") that names the preferred URL for a page when the same or similar content is reachable at more than one address. A correct canonical tag helps crawlers consolidate duplicate or near-duplicate URLs onto one indexed version instead of splitting signals across several.
Checkout eligibility Structured data
A product-level attribute or flag in a merchant feed declaring that an item may be bought inside an AI product rather than only described there. Google states that only listings carrying its native_commerce checkout_eligibility attribute display the Buy button for its checkout experience, and OpenAI states that setting an eligibility flag does not by itself complete checkout onboarding. The attribute expresses intent; the integration is separate work.
Claim-and-source record Content strategy & quality
A record that pairs each material factual claim in a piece of content with the specific source that supports it and any limitation on that evidence. Building one before publication makes unsupported assertions and citation gaps visible to an editor instead of surfacing only after a reader or AI system questions the claim.
Client-side rendering (CSR) Crawlability
A rendering approach where the browser downloads mostly empty HTML and JavaScript builds the visible content after the page loads. A crawler that does not execute JavaScript, or limits how much it executes, can miss content that only client-side rendering produces.
Context window
The limited amount of instructions, conversation, retrieved passages, tool results, and generated output a language model can reference for one response. A context window is temporary working material, not the model's training dataset; if useful evidence is omitted, truncated, or overwhelmed by other material, the answer may miss it.
Correction request Auditing & correcting
A documented ask, sent to a publisher, platform, or AI provider, to fix a false, outdated, or misattributed statement about an entity. A correction request should point to the specific answer or source, state the accurate fact, and reference supporting evidence, because an unspecific complaint is harder to act on.
Crawl budget Crawlability
The number of pages, and the rate, a given crawler is willing or able to request from a site within a period. A limited crawl budget means low-value or duplicate URLs can consume requests that would otherwise reach important pages, so crawl efficiency affects how completely a site gets seen.
Crawlability Crawlability
The ability of an intended automated system to discover and request a public URL under the site's technical and policy controls. A crawlable page can still fail during delivery, rendering, indexing, retrieval, or answer selection, so crawlability is one stage rather than proof of visibility.
Crawler Crawlability
An automated program that requests and downloads pages to discover, evaluate, or store their content, operated by a search engine, an AI company, an archive, or another service. Search crawlers, AI crawlers, and other bots can behave differently, so treating them as one undifferentiated group can hide which one actually needs access.
D
Disambiguation Entity identity & authority
The process of establishing which specific real-world entity a name, page, or record refers to, when the same or a similar name could apply to more than one thing. Weak disambiguation lets systems attribute facts, reviews, or coverage to the wrong entity, such as merging two businesses with a similar name.
E
E-E-A-T Content strategy & quality
Experience, Expertise, Authoritativeness, and Trustworthiness: a framework Google's quality raters use to describe qualities of credible content and its creators. E-E-A-T is not a ranking factor with a measurable score; it describes the kind of demonstrated experience and sourcing that content should show for a reader to trust it.
Editorial review Content strategy & quality
The process of checking a piece of content's claims, sources, reasoning, and wording against evidence and quality standards before it is published or after it is flagged. Editorial review assigns a responsible person to approve the result, distinguishing it from an automated check or a spell-check pass.
Embedding
A numerical representation of text, an image, or another item that lets a system compare relatedness. Retrieval systems can compare a query embedding with stored passage embeddings to find semantically similar material, including material that uses different words; similarity alone does not establish factual support or authority.
Entity Entity identity & authority
A distinct, identifiable thing, such as an organization, person, product, or place, that a system can describe, reference, and connect to facts and relationships. Search and AI systems increasingly reason about entities and their attributes rather than matching keywords alone.
Entity authority Entity identity & authority
This publication’s working term for the credible evidence associated with a particular entity and subject. Clear identity helps distinguish the entity; independent recognition and well-supported expertise are separate evidence questions. Consistent profiles alone do not establish authority, and there is no universal entity-authority score.
F
Fluency optimization Content strategy & quality
A named rewriting tactic in the 2024 GEO study that asked a language model to improve the fluency of website text. The paper does not define a universal reading-grade target or platform rule. On this site, the durable practice is clear, accurate prose that preserves the subject, conditions, and evidence when a passage is read separately; it does not guarantee retrieval or citation.
G
GEO-bench Measurement
The 10,000-query benchmark introduced in the 2024 GEO study to evaluate proposed generative-engine optimization methods. It combines cleaned text from the top five Google results with diverse query sources and a simulated two-step retrieval-and-generation setup. Its scores describe that research setting; they are not a report of a website's organic performance on a current answer product.
Generative engine optimization (GEO)
Practices aimed at improving how generative AI systems represent, summarize, or recommend a brand, product, or piece of content in their output. Generative engine optimization overlaps heavily with answer engine optimization; this site treats AI SEO as the umbrella term for both.
Generative-engine impression Measurement
A paper-specific way of describing a source's visibility within one generated response, rather than its rank in a list of links. The GEO study measures the cited source's relative word count, position, and LLM-judged presentation. It is useful for interpreting that study, but it is not a standardized platform metric and should not be confused with Search Console impressions, traffic, or business outcomes.
Ground truth Auditing & correcting
The best verified reference evidence for a specific claim, time, market, and version, used to assess an answer. The appropriate source depends on the claim: a maintained internal record may establish a product price, while an independent finding requires independent evidence. Record disagreements and uncertainty rather than treating an organization’s preferred account as truth.
Grounding
Connecting generation to supplied evidence, such as retrieved documents or a knowledge base. The system still has to select relevant material and use it faithfully; the presence of grounding does not establish factual accuracy or guarantee visible citations.
Grounding query Measurement
An informal AI-search term for a query a system uses to retrieve supporting material before composing an answer. Its exact meaning is product-specific and usually not observable. In Bing Webmaster Tools AI Performance, a grounding query is instead a short, aggregated reporting phrase that groups citation activity around a recurring retrieval theme. That Bing report is not an individual user prompt, a complete query log, a ranking, or a causal explanation for a citation; one grounding query can map to several pages, and one page can appear under several grounding queries.
GTIN Structured data
A Global Trade Item Number, the standardized product identifier that appears as a UPC, EAN, or ISBN depending on the region and item type. Merchant programs use it to match the same product across catalogs, so a GTIN that differs between a feed, a product page, and its structured data breaks that match. Spreadsheet exports are a common cause, because they strip leading zeros unless the column is stored as text.
H
Hallucination
Plausible-sounding AI output that is factually incorrect, unsupported, or invented. Hallucination happens because a language model generates a statistically likely continuation of text rather than verifying each claim against a source, which is one reason AI answers need an audit and correction process.
HTTP status code Crawlability
The three-digit code an origin server returns with each response, such as 200 for success, 301 for a permanent redirect, 404 for a missing page, or 410 for a page removed on purpose. Crawlers use these codes to decide whether to index, drop, or transfer ranking signals for a URL, so an incorrect code can hide or misdirect a page.
Human-led content Content strategy & quality
Content where a person remains accountable for the question addressed, the sources used, the reasoning, the final wording, and later corrections, even when AI tools assist with research, drafting, or editing. Human-led does not mean every sentence is typed without assistance; it means editorial judgment and responsibility are not delegated to automation.
Hydration Crawlability
The step where JavaScript attaches interactivity to HTML that was already rendered, either on the server or as a static file, turning static markup into a fully working page in the browser. Content present before hydration is visible to more crawlers than content that only appears afterward, which is why pre-hydration HTML deserves a direct check.
I
Indexability Crawlability
A processed page's eligibility to be stored and considered as a particular URL by an index. Indexability depends on controls and signals such as noindex directives, response status, canonicals, and duplicates; it does not guarantee that a system will index, rank, retrieve, or cite the page.
Indexing Crawlability
The process by which a search engine or AI system stores and organizes a crawled page so it can be retrieved later for a relevant query. A page can be successfully crawled and rendered and still be excluded from an index, so crawling, rendering, and indexing should be checked as separate steps rather than assumed to succeed together.
J
JSON-LD Structured data
JSON for Linked Data: a JSON-based format for embedding structured data in a page, usually inside a single script type="application/ld+json" element. JSON-LD is the format most commonly recommended for implementing Schema.org markup because it can describe a page's entities without being woven into the visible HTML.
K
Knowledge Graph Entity identity & authority
A structured database of entities and the relationships between them, used by a search engine or AI system to resolve identity and assemble facts about a subject. A knowledge graph entry for an organization or person can draw on structured data, sameAs links, and other corroborating sources rather than any single page.
L
Large language model (LLM)
A machine learning model trained on large volumes of text to predict and generate language, underlying most AI chatbots and generative search features. A large language model's output reflects patterns in its training data and any retrieved context, not a lookup against a maintained, current source of truth by default.
Light DOM Crawlability
The ordinary child nodes written inside a web component's host element, outside its shadow tree. A component can display those nodes directly or assign them to slots in a shadow DOM, so original HTML, the document tree, and the composed browser result may contain different structures.
llms.txt Crawlability
A proposed plain-text file, placed at a website's root, intended to give AI systems a curated summary of a site's most important content and links. Unlike robots.txt, llms.txt is not a web standard that major crawlers are confirmed to read and act on, so it should supplement, not replace, established access and structured-data practices.
Log file analysis Crawlability
Reviewing a server's raw access logs to see which crawlers actually requested which URLs, how often, and with what response codes. Log file analysis shows verified crawler behavior, which can differ from what a sitemap, robots.txt file, or crawl-simulation tool predicts.
M
Model weights
The numerical parameters adjusted during model training that encode learned language patterns and behavior. Model weights are not a live, browsable source library: supplying current documents through retrieval or a prompt can influence one response without rewriting the deployed model's weights.
N
NAP Entity identity & authority
Name, address, and phone number: identity and contact details for a local business. Check that records identify the correct location and give accurate current information. Harmless punctuation or formatting differences are different from conflicting addresses, names, or phone numbers.
Noindex Crawlability
A directive, delivered as a robots meta tag or an X-Robots-Tag HTTP header, that asks a crawler not to include a page in its index. A page can carry noindex and still be crawled, so blocking a crawler entirely and asking it not to index a page are different actions with different effects.
P
Product feed Structured data
A file a merchant supplies to a specific company describing its catalog, with one row or object per item or variant, carrying identifiers, title, description, URL, image, price, and availability. Unlike structured data, which describes a page a system has already fetched, a feed is pushed on a schedule and covers products regardless of whether their pages were crawled. Google's product data specification is the common baseline that most AI merchant programs build on.
Prompt
The instruction, question, or input a user or system gives to a language model to produce a response. The exact wording, context, and any retrieved documents included with a prompt can change the resulting answer, which is one reason the same question can produce different AI answers on different attempts.
Prompt set Measurement
A versioned collection of questions used to observe AI answers under defined conditions. A defensible prompt set is drawn from a documented audience-question frame, preserves important groups and edge cases, and separates a stable comparison core from exploratory prompts instead of selecting only convenient examples.
Position-adjusted word count Measurement
An impression metric proposed in the GEO study. It allocates the words in answer sentences associated with a cited source, then gives earlier sentence positions greater weight through an exponential decay. It is an experimental proxy for citation prominence in that study's responses, not a count supplied by answer engines or evidence of reader attention, clicks, or conversions.
R
Referral visit Measurement
A visit that reaches a website through a link on another surface, including an AI answer. Analytics can identify an AI source only when the available referrer or campaign information and collection rules support that classification. Missing information, consent, and redirects can limit attribution.
Rendering Crawlability
The process of turning a page's HTML, CSS, and JavaScript into the content a user or crawler ultimately sees, which can happen on the server before delivery, in the browser after delivery, or in stages across both. What a crawler indexes depends on what it renders, not only on what the initial HTML response contains.
Retrieval-augmented generation (RAG)
A method that retrieves documents or passages and supplies them as context for generation. Retrieved material can include information outside the model’s training data, but it may be incomplete, stale, or poorly matched. Retrieval, faithful use of evidence, and displayed citations require separate checks.
Rich result Structured data
An enhanced search result display, such as a review star rating or a recipe card, that a search engine can generate when a page's structured data matches its eligibility requirements. Valid structured data makes a page eligible for a rich result; it does not guarantee the feature will appear.
Robots meta tag Crawlability
An HTML meta name="robots" element, or an equivalent HTTP header, placed on a page to give per-page instructions to crawlers, such as noindex or nofollow. Robots meta tag instructions apply to that individual page and are separate from the site-wide rules set in robots.txt.
robots.txt Crawlability
A plain-text file at a website's root that tells crawlers which parts of the site they may or may not request, addressed to one or more named user agents. robots.txt controls crawling access; it does not remove a page from an index by itself and does not guarantee that a named crawler will comply.
S
sameAs Entity identity & authority
A Schema.org property that links an entity's structured data to its verified profiles on other websites, such as an official social media account or a reference database entry. sameAs gives corroborating identity signals that can help a system confirm which real-world entity a page describes.
Sampling frame Measurement
The documented list or selection procedure that defines which audience questions are eligible for a prompt sample. Its sources, scope, exclusions, markets, languages, and time period limit what conclusions the resulting observations can support.
Schema.org Structured data
A shared vocabulary of types and properties, maintained by a multi-company collaboration, used to describe entities and content in structured data. Schema.org defines the terms available; JSON-LD is the format most often used to express them on a page.
Server-side rendering (SSR) Crawlability
A rendering approach where the server builds the page's HTML, including its main content, before sending it to the browser or crawler. Content present in server-side-rendered HTML is visible to a wider range of crawlers than content that depends on client-side JavaScript execution.
Shadow DOM Crawlability
A DOM tree attached to a host element to encapsulate a web component's internal structure and styles. Open and closed shadow roots expose different inspection access, and ordinary serialization of the document element does not include shadow-tree content, so rendering checks must inspect web components explicitly.
Share of voice Measurement
The share of relevant AI answers, mentions, or citations that name a given brand or source compared with its competitors, for a defined set of prompts or queries. Share of voice is a comparative measurement; it depends entirely on which prompts, timeframe, and competitor set were sampled.
Stratified sampling Measurement
A sampling method that divides an eligible population into meaningful groups, or strata, and selects observations within each group. In AI-answer measurement, it can keep a large prompt category from hiding a smaller market, intent, or high-consequence topic that needs its own coverage.
Structured data Structured data
Machine-readable information about a page's visible content, most often implemented with Schema.org vocabulary in a JSON-LD block. Structured data describes what is already on the page; it is not a synonym for content quality or a way to add facts a page does not otherwise support.
Subjective impression Measurement
An LLM-judged GEO-study metric for a citation's perceived relevance, influence, uniqueness, position, amount of associated material, click likelihood, and diversity. The study normalized these judgments for comparison with its position-adjusted word-count metric. It is a research evaluation construct, not a reliable direct measure of human attention or a standard answer-engine report.
T
Training data
The text, code, and other material a language model is trained on to learn statistical patterns of language. A model's knowledge of events, facts, and content published after its training cutoff is limited unless it uses retrieval or another live-data mechanism.
U
Universal Commerce Protocol (UCP) Structured data
An open standard for agentic commerce that extends existing merchant product data so it can support a transaction rather than only a listing. Google describes it as an evolving standard and notes that not all features in the specification are available on its surfaces; Microsoft asks merchants for a UCP-compliant feed in Microsoft Merchant Center. For a merchant already maintaining a Merchant Center feed, UCP is mostly additional attributes rather than a new file.
User agent Crawlability
A string a crawler or browser sends with each request identifying itself, such as its name and sometimes its version. robots.txt rules and server logs both rely on the user agent string to distinguish one crawler, including one AI crawler, from another.
V
Vector search
A retrieval method that compares numerical representations to find items that are semantically related to a query, even when they share few exact keywords. A vector-search result is a candidate for later ranking or generation; similarity does not prove that the passage is accurate, current, authoritative, or used in the final answer.
X
XML sitemap Crawlability
An XML file that lists a site's URLs, and optionally their last-modified dates, to help crawlers discover pages efficiently. An XML sitemap is a discovery aid: listing a URL in a sitemap does not guarantee that it will be crawled, rendered, or indexed.
Z
Zero-click search Measurement
A search or AI interaction where the user gets their answer directly on the results or answer surface and does not click through to any website. A rising share of zero-click searches means citation and mention tracking matter alongside referral-visit counts, since a page can be used in an answer without ever producing a click.