Resources / Crawlability, Rendering & Indexability / Checklist

Crawlability, rendering, and indexability checklist

Review crawler policy, server delivery, discovery, rendering, index controls, and verification across 29 technical checks.

What this checklist covers

Use this checklist to review whether search crawlers and AI agents can find your public pages, retrieve their content, and process it as intended. It begins with crawler policy and server access, then covers discovery, rendering, index controls, optional formats, and ongoing verification.

For a new website, use the checklist during development and repeat the checks on the live site. For an existing website, review representative pages and templates, then repeat affected checks after changes. Record a finding, an owner, and a follow-up for each applicable item.

Access for model training, automatic search crawling, and a page fetch requested by a user are separate decisions. The recommendations below distinguish those purposes; passing these checks does not guarantee indexing, inclusion in an AI answer, or a citation.

Read Crawlability, Rendering & Indexability first if you need the concepts, system boundaries, failure patterns, or an assessment workflow before working through the individual checks.

What to do first

Choose the public pages and systems you intend to support. If an important page is blocked, returns an error, or omits essential content for an intended system, investigate that demonstrated failure first. Preserve the response, identify the failing stage, and give the responsible owner a checkable repair.

Next: Recheck the affected response and content, then follow the relevant discovery, rendering, or indexing checks for your objective. Checks that already pass do not require a new implementation.

Can wait: Additional AI-oriented files or alternate formats without an identified consumer or use case. Component-specific inspection applies when your implementation or evidence calls for it.

Move on when: The affected pages pass the relevant checks and the evidence and recheck trigger are recorded. Passing an access check does not establish indexing or citation.

Jump to the relevant module

Use a working register

Copy this compact register into the audit record. One row can represent one affected template or URL group; add a separate row when the evidence or owner differs.

CheckURL or templateResultEvidence savedOwnerNext action and recheck trigger
8. HTTP responsesProduct detailNeeds workRedirect chain and response bodyPlatformRepair loop; retest after release
17. RenderingProduct detailPassOriginal and rendered capturesWeb platformRecheck after product-component change
24. Index directivesArchiveNot applicablePolicy recordContent operationsRecheck if archive becomes public

Use pass, needs work, or not applicable only after recording the evidence that supports it. “Not applicable” means the condition does not exist for that template, not that it was skipped.

Crawler access and content use

1. Review robots.txt rules

Check that /robots.txt is reachable and that its rules allow the crawlers, pages, and supporting files you intend to make accessible. The file gives instructions to cooperating crawlers; it is not a password system, and blocking a URL there does not necessarily remove that URL from search results.

Use the robots.txt guide and copyable policies to separate conventional search, AI search, training, and user-requested access decisions.

2. Decide whether to allow training crawlers

Record whether your organization permits collection for model training, then review the provider-specific controls that express that choice. OpenAI identifies GPTBot and Anthropic identifies ClaudeBot with training-related collection, distinct from their search and user-requested agents. Do not treat a training opt-out as a decision to block all AI access.

The robots.txt policy templates show how to block named training controls without also blocking search crawlers.

3. Review search and AI search crawler access

Check access for the search services you want to reach, including conventional search crawlers and AI search crawlers such as OAI-SearchBot and Claude-SearchBot. These agents collect information for search functions, which is a separate purpose from fetching a page at the moment a user asks about it.

Use the robots.txt guide to keep named search access separate from training policy.

4. Review user-requested web fetches

Check what happens when an assistant requests a page to help answer a user’s question, including whether it receives the content or a challenge page. Provider rules differ: OpenAI says robots.txt rules may not apply to ChatGPT-User requests, while Anthropic says its bots, including Claude-User, honor those directives. Review each provider’s policy rather than assuming one rule covers every agent.

The robots.txt guide explains what those provider differences mean for an access policy.

5. Check provider-specific content-use controls

Identify controls that govern how already-crawled information may be used, as well as controls that prevent fetching it. Google’s Google-Extended token covers specified Gemini training and grounding uses, but is not a separate requesting crawler and does not control inclusion in Google Search. Record the products and uses affected before changing a setting.

See the robots.txt guide for a copyable Google-Extended example alongside OpenAI and Anthropic controls.

6. Verify crawler identity

Use a provider’s documented verification method, such as published IP ranges or forward and reverse DNS checks, before granting special access to a claimed crawler. A user-agent name alone is not sufficient evidence of identity; use the available verification information alongside your request logs.

Server access and delivery

7. Check domain, HTTPS, and connection reliability

Verify that public URLs resolve to the intended server, establish a valid HTTPS connection, and respond consistently from outside your own network. Include relevant subdomains and asset hosts in the review so a working homepage does not hide a connection failure elsewhere.

8. Check HTTP responses and redirects

Review whether live pages return successful responses, moved pages lead to the correct destination, and missing pages return an appropriate error. Look for redirect loops and pages that report success while displaying an error, an empty shell, or the wrong content. A successful response is a delivery check, not proof of indexing.

9. Review firewalls, bot protection, and challenges

Check whether your hosting platform, content delivery network, or web application firewall blocks intended crawlers or substitutes a CAPTCHA or browser challenge for the page. Reconcile these settings with your crawler policy, using narrow, verified exceptions where appropriate rather than disabling protection across the site.

10. Check access without a login or saved session

Test public pages without your own account, cookies, or stored location preferences, and inspect the content actually returned. Identify consent screens, account prompts, or other gates that replace intended public information, while retaining authentication for material that should remain private.

11. Review response time, page size, and rate limits

Check slow responses, oversized pages, timeouts, and repeated rate-limit or server errors in crawler requests. Keep the site able to serve legitimate traffic within its capacity, and investigate the cause of errors before changing limits; different fetchers have different processing constraints.

12. Check caching and content freshness

After an update, confirm that the public page and any alternate formats serve the current information rather than an outdated cached copy. Review cache invalidation and change signals such as ETag and Last-Modified, which supported crawlers can use when deciding whether a resource has changed.

Page discovery and URL management

Make important pages reachable through ordinary links with usable destination URLs and descriptive link text. Review orphan pages and navigation that depends entirely on scripts or form submissions, since a visible control is not necessarily a link a crawler can follow.

14. Review XML sitemaps

Check that XML sitemaps are accessible, valid, and list the preferred URLs you want discovered, with accurate modification dates where supplied. Keep sitemap indexes and search-engine submissions current as the site changes. A sitemap supports discovery but does not guarantee crawling or indexing.

Follow the XML sitemap tutorial to generate, validate, submit, and maintain the files.

15. Check canonical URLs and duplicate versions

Review the preferred URL declared for each page and check that internal links, redirects, and sitemap entries support that choice. Look for conflicting versions created by alternate hosts, tracking parameters, or duplicate paths; a canonical declaration is a signal to search engines, not an access restriction.

16. Review filters, pagination, and crawl traps

Check whether filters, sorting options, calendars, and paginated listings create excessive duplicate URLs or leave important items unreachable. Decide which combinations deserve discovery and provide a finite, navigable route to the underlying pages. Apply restrictions carefully so that reducing unnecessary crawling does not hide useful content.

Content and rendering

17. Compare original, rendered, light DOM, and shadow DOM content

Compare the initial HTML response, before JavaScript runs, with the rendered document after scripts have executed. For web components, inspect light DOM content, open shadow roots, slot assignment, and the composed result a user receives instead of relying on outerHTML, which does not serialize shadow-tree contents. Check the main text, links, metadata, and structured data for omissions or changes, and record which content depends on rendering.

Google says its renderer flattens light DOM and shadow DOM into rendered HTML, but that Google-specific behavior does not establish how another search crawler, AI fetcher, or browser-using agent processes the component. A closed shadow root restricts ordinary page-script inspection; its mode alone does not prove whether a system can render, index, or understand the visible content.

Use the original-versus-rendered HTML playbook for a repeatable capture, comparison, and recheck procedure.

18. Check JavaScript, styles, and data dependencies

Verify that the scripts, stylesheets, and data requests needed to produce the page are accessible and complete without errors. A reachable page can still lose important content when a required resource is blocked or fails, so review those dependencies alongside the page itself.

The rendering comparison playbook includes dependency-failure tests and a record for each mismatch.

19. Check content that requires interaction

Review information loaded only after scrolling, clicking, expanding a control, or using an internal search form. Google Search does not interact with pages to trigger that loading; check that essential information has an accessible page or loads through a supported method rather than assuming a crawler will perform the interaction.

Use the rendering comparison playbook to test the initial load before interaction and record what remains unavailable.

20. Review semantic HTML and accessible controls

Use meaningful headings, landmarks, links, and labeled controls so the document’s structure and available actions can be understood. Check keyboard access and text alternatives as part of accessibility for people, and test any browser-using AI agent separately rather than treating human accessibility compliance as proof of agent compatibility.

21. Check documents, images, and other media

Review important PDFs, images, and other files for accessible URLs, appropriate content types, and any indexing restrictions sent in response headers. Provide readable text for essential information rather than assuming every system will extract it from a scan, image, or video; supported file types differ by service.

22. Compare mobile and desktop content

Check that mobile and desktop versions provide equivalent essential content, links, and indexing instructions. Google uses the mobile version for indexing, so a desktop-only review can miss omissions in the content Google processes.

23. Check language and regional versions

Give intended language and regional versions discoverable URLs, links between alternatives, and consistent language annotations where applicable. Review location-based redirects and cookie-dependent content so visitors and crawlers can reach the intended version without being forced into another one.

Indexing and answer display

24. Review meta robots and X-Robots-Tag

Check page-level robots tags and HTTP response headers for unintended noindex or other restrictions, including settings inherited from development environments. A crawler must be able to retrieve the response to read these instructions, so blocking the same URL in robots.txt can prevent it from seeing the rule.

25. Review snippet and preview controls

Check whether nosnippet, max-snippet, or data-nosnippet settings restrict how your content can appear in supported search features. Google requires a page to be indexed and eligible for a search snippet to appear as a supporting link in AI Overviews or AI Mode. Keep these display controls separate from your decisions about model training.

Optional formats for AI readers

26. Evaluate llms.txt

llms.txt is a proposed Markdown file that gives agents a concise introduction to a site and links to useful material. Treat it as an optional aid for systems that use it, not a substitute for robots.txt, sitemaps, or accessible pages; Google says its AI search features do not require an AI text file. If you publish one, check its accuracy, links, and actual use before attributing any benefit to it.

27. Review Markdown or other alternate content versions

If you provide Markdown pages or another text-oriented format, verify that they are discoverable and carry the same current facts as the main pages. The llms.txt proposal describes Markdown alternatives, but this is an optional delivery choice rather than a universal requirement. Apply the intended access policy to every version and avoid creating an unmaintained second copy.

Verification and maintenance

28. Inspect crawl and fetch logs

Review server and security logs to see which verified agents requested which URLs, what responses they received, and where failures repeat. A logged request shows access activity; it does not by itself show that the content was indexed, used for training, or cited in an answer.

29. Check search-engine reports and repeat tests after changes

Use search-engine inspection and indexing reports to compare live accessibility with the version a search engine has processed. Test representative templates after launches, migrations, or changes to hosting, security rules, rendering, and URL structure, retaining the findings and correction owner. Google’s live URL test can help diagnose access but does not guarantee indexing or represent every AI agent.