Resources / Crawlability, Rendering & Indexability / Guide
Track AI referrals and control AI crawlers
Maintain a verified register of answer-engine referrals, URL evidence, and documented crawler controls without classifying ordinary search traffic as AI traffic.
Keep referral evidence and crawler policy separate
One company can operate search, chat, training collection, and user-requested fetching. A referral identifies a recorded visit; a robots.txt group expresses a crawl-access policy. Neither proves that the other happened. Maintain separate registers and update both only from observed traffic or provider documentation.
Do not classify google.com as an AI referral: it includes ordinary Google Search alongside AI features. Do not classify bare bing.com as Copilot traffic for the same reason. bing.com/chat is a historical Bing Chat path; the current Copilot surface uses copilot.microsoft.com. Preserve an observed bing.com/chat pattern for historical reporting, but add copilot.microsoft.com only when it appears in the traffic you collect. Cross-origin browsers often send only an origin, so a GA4 source of bing.com alone is unknown, not an AI-answer referral.
Referral register: candidates to inspect
Start with raw session source/medium, server Referer, landing-page query string, and any approved campaign field. Record the exact observed value, date, sample size, owner, and exclusions before creating a channel rule. Do not use a model name as an analytics source: models do not normally send referrals; products and links do.
Referral patterns are a moving target. Before setting up or revising a segment, filter, or channel rule, inspect a sufficiently long history of referring URLs for the property. Identify the exact domains and, where the data retains them, the relevant subfolders or paths. This historical inventory is the basis for the rule; a public list is only a set of candidates to investigate. Version the rule when a provider changes domains or paths, and keep older verified patterns when they are needed to interpret past reporting.
| Product or surface | Candidate evidence to inspect | Classification rule |
|---|---|---|
| ChatGPT | chatgpt.com, chat.openai.com, or an observed approved URL parameter | Match only exact, observed values. |
| Claude | claude.ai | Match only after it appears in raw traffic. |
| Perplexity | perplexity.ai | Match only after it appears in raw traffic. |
| Gemini | gemini.google.com | Do not widen to google.com. |
| Microsoft Copilot | copilot.microsoft.com; a preserved bing.com/chat path | Never treat bare bing.com as Copilot. |
| Grok | grok.com or an observed Grok-specific path | Do not group all x.com traffic. |
| DeepSeek | chat.deepseek.com | Match only after it appears in raw traffic. Do not infer it from deepseek.com marketing or API traffic. |
| Poe | poe.com | Match only after it appears in raw traffic; Poe hosts many underlying models, so do not relabel the visit as the model provider. |
| Meta AI (including Meta’s Muse Spark model context) | meta.ai or ai.meta.com | Match only an observed dedicated Meta AI origin. Do not group Facebook, Instagram, Messenger, WhatsApp, or Threads traffic. Model names such as Muse or Muse Spark are not referral domains or documented robots.txt tokens. |
| Mistral Le Chat | chat.mistral.ai | Match only after it appears in raw traffic. |
| Qwen Chat | chat.qwen.ai | Match only after it appears in raw traffic. |
| Kimi, Phind, You.com | kimi.com, phind.com, you.com | Match only the observed product origin; do not broaden to unrelated partner or model domains. |
| Kagi Assistant, Duck.ai, Brave Leo | a preserved assistant-specific path or exact observed origin | Their parent domains also serve conventional search or browsing. A bare parent domain is unknown until path-level or other arrival evidence establishes the surface. |
UTM parameters are campaign declarations, not proof of an organic answer referral. Preserve them separately: utm_source=chatgpt.com can identify a deliberately tagged link, while an untagged chatgpt.com referrer is a different arrival mechanism. Never manufacture UTMs to validate an organic rule.
Crawler-control register
These are documented controls worth reviewing. Allow means omit a named block and let the applicable general policy govern it; do not add a named Allow: / group unless you repeat every restriction that should remain in force.
| Provider use | Token or agent | What a decision controls |
|---|---|---|
| OpenAI training | GPTBot | Training-related collection. |
| ChatGPT search | OAI-SearchBot | ChatGPT search surfacing. |
| OpenAI user fetch | ChatGPT-User | User-triggered fetch; OpenAI says robots rules may not apply. |
| Anthropic training | ClaudeBot | Training-related collection. |
| Claude search | Claude-SearchBot | Search-related collection. |
| Claude user fetch | Claude-User | User-triggered fetch; Anthropic says its bots honor robots.txt. |
| Perplexity search | PerplexityBot | Perplexity search crawling. |
| Perplexity user fetch | Perplexity-User | User-triggered fetching; Perplexity documents different robots behavior. |
| Google Gemini use | Google-Extended | Gemini training and specified grounding uses; it is a control token, not a separate request user agent, and does not affect Google Search ranking or inclusion. |
| Bing and Copilot grounding | bingbot | Bing’s documented standard crawler; blocking it can affect Bing indexing and related Copilot grounding. Treat it as a Bing search decision, not a Copilot-only toggle. |
For any other claimed AI crawler—Amazon, Meta, Mistral, ByteDance, Qwen, or a new vendor—add it only after locating the vendor’s current crawler documentation. A string found in an unverified log can be spoofed.
Use a decision record before editing robots.txt
For each token, record purpose, allowed paths, business rationale, provider documentation URL, effective date, owner, and recheck date. Keep private data behind authentication; robots.txt is public and voluntary. Test the live file and representative paths after every change, then review verified logs. The robots.txt guide supplies safe policy patterns; the GA4 referral tutorial supplies channel-rule validation.