Resources / Crawlability, Rendering & Indexability / Guide

robots.txt for Search and AI Crawlers

Copyable robots.txt policies for conventional search, AI search, model training, and public paths—plus the checks that keep them from conflicting.

What this file controls

robots.txt tells compliant automated clients which URLs they may request. Put the plain-text file at the root of each host or subdomain it governs, such as https://www.example.com/robots.txt. Rules on www.example.com do not automatically govern shop.example.com.

The file manages crawling, not confidentiality or search removal. A blocked URL can still be discovered from links and appear in search without a useful snippet. Protect private material with authentication. If a public page should be excluded from an index, use a supported noindex directive and allow the crawler to retrieve it.

Make four decisions before editing

Write down the intended policy for each purpose:

  1. Conventional search crawling and indexing.
  2. Automated crawling for AI search or answer retrieval.
  3. Collection that may be used for model training.
  4. A fetch initiated by a user inside an AI product.

Providers use different controls for these purposes. OpenAI documents OAI-SearchBot for ChatGPT search and GPTBot for content that may be used in model training. Anthropic documents Claude-SearchBot, ClaudeBot, and Claude-User separately. Google documents Google-Extended as a control token for specified Gemini training and grounding uses; it is not a separate HTTP crawler, and its rules do not affect Google Search inclusion or ranking.

OpenAI says robots.txt rules may not apply to ChatGPT-User because its requests are initiated by a person. Anthropic says its bots, including Claude-User, honor robots.txt. A rule therefore expresses a request to the clients that support it, not a universal access control.

Template: allow search and block named training uses

This common starting point leaves conventional search and documented AI search crawlers under the general rules while blocking the named training controls. Replace the example paths and domain before publishing.

User-agent: *
Disallow: /account/
Disallow: /checkout/
Disallow: /internal-search/

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml

Do not add a named OAI-SearchBot or Claude-SearchBot group merely to say Allow: /. Under Google’s interpretation, a specific matching group can take precedence over the * group, so shared private-path restrictions would need to be repeated inside it. Omitting those named groups lets compliant search bots use the general policy.

Read the policy as a result, not a bot list

In the example above, a compliant crawler matched only by User-agent: * may request /guides/, but not /account/ or /internal-search/. GPTBot, ClaudeBot, and Google-Extended match their named groups and may request no path at all. The absent OAI-SearchBot group means that crawler falls back to the wildcard restrictions. If you add a named group for it, repeat every restriction that should still apply:

User-agent: OAI-SearchBot
Disallow: /account/
Disallow: /checkout/
Disallow: /internal-search/

This is why a named Allow: / can accidentally open paths that the wildcard group blocked. Test the policy against representative public and private paths before publishing it.

Template: block named automated AI crawling

Use this only after deciding that the loss of AI search visibility is acceptable. It blocks the named OpenAI and Anthropic search crawlers as well as their training crawlers. It cannot prevent every user-initiated fetch, and it does not block conventional search crawlers unless another rule does so.

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml

Add a separate Claude-User block if your policy also rejects user-requested Claude fetches:

User-agent: Claude-User
Disallow: /

OpenAI directs publishers to use OAI-SearchBot for search opt-outs and notes that ChatGPT-User rules may not apply. Treat network enforcement, authentication, or application authorization as separate controls when access must be prevented.

Template: keep public crawling away from private or low-value paths

This policy applies the same path restrictions to every compliant client that does not have a more specific matching group.

User-agent: *
Disallow: /account/
Disallow: /checkout/
Disallow: /preview/
Disallow: /internal-search/

Sitemap: https://www.example.com/sitemap.xml

Do not list secrets in robots.txt; the file is public. Use authentication for private pages. Also avoid blocking CSS, JavaScript, images, or data endpoints needed to render public pages unless you have verified the result for every crawler you intend to support.

Implement the policy

  1. Inventory the live file, CDN rules, firewall rules, meta robots tags, and X-Robots-Tag headers. These layers can conflict.
  2. Replace every example domain and path. Preserve any existing directives you still need.
  3. Use the exact user-agent token published by the provider. Do not infer a token from a browser user-agent string.
  4. Keep each named agent’s rules together. If the same agent has multiple matching groups, parsers may combine them.
  5. Upload the file as UTF-8 plain text at /robots.txt on each applicable host.
  6. Request the live URL without a logged-in session, then test representative allowed and disallowed paths.
  7. Review request logs after deployment. A user-agent string can be spoofed, so use provider-published IP data or DNS verification before creating privileged firewall exceptions.

Validate the result

Run these checks against the public host:

curl -i https://www.example.com/robots.txt
curl -L https://www.example.com/robots.txt

Confirm a successful response, a plain-text body, the intended host, and the absence of login pages, redirects to another environment, or injected HTML. Then inspect important public pages with the relevant search-engine tools and review server logs for repeated 401, 403, 429, or 5xx responses.

Check the policy again after a migration, CDN or firewall change, subdomain launch, or provider documentation change. A syntactically valid file can still encode the wrong business decision.

Common failures

Continue the technical review

Use the XML sitemap tutorial to make preferred public URLs discoverable. Then use the raw-versus-rendered HTML playbook to confirm what a fetcher receives before and after JavaScript runs.

Official references