Resources / Crawlability, Rendering & Indexability / Guide

Track AI referrals and control AI crawlers

Maintain a verified register of answer-engine referrals, URL evidence, and documented crawler controls without classifying ordinary search traffic as AI traffic.

Keep referral evidence and crawler policy separate

One company can operate search, chat, training collection, and user-requested fetching. A referral identifies a recorded visit; a robots.txt group expresses a crawl-access policy. Neither proves that the other happened. Maintain separate registers and update both only from observed traffic or provider documentation.

Do not classify google.com as an AI referral: it includes ordinary Google Search alongside AI features. Do not classify bare bing.com as Copilot traffic for the same reason. bing.com/chat is a historical Bing Chat path; the current Copilot surface uses copilot.microsoft.com. Preserve an observed bing.com/chat pattern for historical reporting, but add copilot.microsoft.com only when it appears in the traffic you collect. Cross-origin browsers often send only an origin, so a GA4 source of bing.com alone is unknown, not an AI-answer referral.

Referral register: candidates to inspect

Start with raw session source/medium, server Referer, landing-page query string, and any approved campaign field. Record the exact observed value, date, sample size, owner, and exclusions before creating a channel rule. Do not use a model name as an analytics source: models do not normally send referrals; products and links do.

Referral patterns are a moving target. Before setting up or revising a segment, filter, or channel rule, inspect a sufficiently long history of referring URLs for the property. Identify the exact domains and, where the data retains them, the relevant subfolders or paths. This historical inventory is the basis for the rule; a public list is only a set of candidates to investigate. Version the rule when a provider changes domains or paths, and keep older verified patterns when they are needed to interpret past reporting.

Product or surfaceCandidate evidence to inspectClassification rule
ChatGPTchatgpt.com, chat.openai.com, or an observed approved URL parameterMatch only exact, observed values.
Claudeclaude.aiMatch only after it appears in raw traffic.
Perplexityperplexity.aiMatch only after it appears in raw traffic.
Geminigemini.google.comDo not widen to google.com.
Microsoft Copilotcopilot.microsoft.com; a preserved bing.com/chat pathNever treat bare bing.com as Copilot.
Grokgrok.com or an observed Grok-specific pathDo not group all x.com traffic.
DeepSeekchat.deepseek.comMatch only after it appears in raw traffic. Do not infer it from deepseek.com marketing or API traffic.
Poepoe.comMatch only after it appears in raw traffic; Poe hosts many underlying models, so do not relabel the visit as the model provider.
Meta AI (including Meta’s Muse Spark model context)meta.ai or ai.meta.comMatch only an observed dedicated Meta AI origin. Do not group Facebook, Instagram, Messenger, WhatsApp, or Threads traffic. Model names such as Muse or Muse Spark are not referral domains or documented robots.txt tokens.
Mistral Le Chatchat.mistral.aiMatch only after it appears in raw traffic.
Qwen Chatchat.qwen.aiMatch only after it appears in raw traffic.
Kimi, Phind, You.comkimi.com, phind.com, you.comMatch only the observed product origin; do not broaden to unrelated partner or model domains.
Kagi Assistant, Duck.ai, Brave Leoa preserved assistant-specific path or exact observed originTheir parent domains also serve conventional search or browsing. A bare parent domain is unknown until path-level or other arrival evidence establishes the surface.

UTM parameters are campaign declarations, not proof of an organic answer referral. Preserve them separately: utm_source=chatgpt.com can identify a deliberately tagged link, while an untagged chatgpt.com referrer is a different arrival mechanism. Never manufacture UTMs to validate an organic rule.

Crawler-control register

These are documented controls worth reviewing. Allow means omit a named block and let the applicable general policy govern it; do not add a named Allow: / group unless you repeat every restriction that should remain in force.

Provider useToken or agentWhat a decision controls
OpenAI trainingGPTBotTraining-related collection.
ChatGPT searchOAI-SearchBotChatGPT search surfacing.
OpenAI user fetchChatGPT-UserUser-triggered fetch; OpenAI says robots rules may not apply.
Anthropic trainingClaudeBotTraining-related collection.
Claude searchClaude-SearchBotSearch-related collection.
Claude user fetchClaude-UserUser-triggered fetch; Anthropic says its bots honor robots.txt.
Perplexity searchPerplexityBotPerplexity search crawling.
Perplexity user fetchPerplexity-UserUser-triggered fetching; Perplexity documents different robots behavior.
Google Gemini useGoogle-ExtendedGemini training and specified grounding uses; it is a control token, not a separate request user agent, and does not affect Google Search ranking or inclusion.
Bing and Copilot groundingbingbotBing’s documented standard crawler; blocking it can affect Bing indexing and related Copilot grounding. Treat it as a Bing search decision, not a Copilot-only toggle.

For any other claimed AI crawler—Amazon, Meta, Mistral, ByteDance, Qwen, or a new vendor—add it only after locating the vendor’s current crawler documentation. A string found in an unverified log can be spoofed.

Use a decision record before editing robots.txt

For each token, record purpose, allowed paths, business rationale, provider documentation URL, effective date, owner, and recheck date. Keep private data behind authentication; robots.txt is public and voluntary. Test the live file and representative paths after every change, then review verified logs. The robots.txt guide supplies safe policy patterns; the GA4 referral tutorial supplies channel-rule validation.

Official references