Topics / Crawlability, Rendering & Indexability

Crawlability, Rendering & Indexability

Important pages can fail before anyone evaluates their content. A system may not discover a URL, may be unable or unwilling to request it, may receive the wrong response, may render an incomplete page, or may decide not to retain the resulting URL. Those are different failures, so they need different evidence and different fixes.

Start with the resource that matches the work in front of you. The technical review checklist records a complete audit. The robots.txt guide turns a crawler-policy decision into rules. The XML sitemap tutorial builds a maintained discovery file. The original-versus-rendered HTML playbook investigates a page whose browser result differs from its delivered response.

This subject matters to AI SEO because an answer system can only use information it can obtain through its own documented access and processing path. Passing a Google test does not establish how another crawler, search product, or user-requested fetcher behaves.

Start with the demonstrated failure

When a material URL has an evidenced access, delivery, rendering, or indexability problem, investigate and recheck that stage before adding an alternate format or expanding discovery signals. If no failure is demonstrated, choose the smallest representative test from the symptom table. The first task is complete when the recorded evidence shows the intended stage working under the relevant conditions; then consider dependent improvements.

On this page

Trace the page through six stages

StageQuestionUseful evidenceIt does not establish
DiscoveryCan a system learn the URL exists?Crawlable links, sitemaps, feeds, external linksThat it requested the URL
AccessIs the intended agent permitted and technically able to request it?robots.txt, authentication and edge policy, verified logsThat the response was usable
DeliveryDid the server return the intended page?Final URL, status, headers, response body, timingThat browser-dependent content appeared
RenderingWhich content, links, and signals exist after processing?Original HTML, rendered DOM, resource failures, inspection outputThat the page will be indexed
IndexabilityIs this version eligible to be retained as the preferred URL?noindex, canonical signals, redirects, duplicate analysisThat it will be selected for a result
Retrieval and useWas it selected for a search result, answer, citation, or visit?Product reports, answer captures, result observation, referralsWhy it was selected

Google documents crawling, rendering, and indexing as distinct stages for JavaScript pages. Its documentation is strong evidence about Google, including its own renderer, but it does not describe every search crawler or AI product. (Google JavaScript SEO basics)

Choose the first test from the symptom

SymptomFirst testNext resource
A crawler should be allowed, but requests are denied or uncertainCheck named-agent policy, identity verification, edge rules, and a verified requestrobots.txt for Search and AI crawlers
Important URLs are missing from discovery pathsCheck internal links, canonical URLs, and sitemap inclusionBuild and validate an XML sitemap
A browser displays content that the initial response lacksSave and compare the response, rendered DOM, links, metadata, and JSON-LDCompare original and rendered HTML
An accessible page is excluded or consolidated elsewhereInspect directives, canonicalization, redirects, duplicates, and the relevant platform reportCrawlability checklist
A citation, mention, or referral is absentCapture the outcome first, then test the earlier stage supported by evidenceWhere to start with an AI SEO problem

Do not use a missing citation as proof of a crawl block, or a successful 200 as proof of indexability. The symptom determines the first layer to inspect.

Separate crawler purpose from crawler access

One provider may operate different agents for conventional search, AI search, training collection, and a fetch triggered by a user. The policy decision is therefore specific: which public material may this named agent request, for which documented purpose?

robots.txt gives instructions to cooperating crawlers. It is not authentication, a way to conceal sensitive data, or a reliable removal mechanism for an already known web-page URL. Google specifically advises using noindex or access control when the aim is to keep a web page out of Google Search. (Google’s robots.txt guide)

Before editing a rule, record the public paths in scope, exact user agent, documented purpose, the provider’s identity-verification method, policy owner, and recheck trigger. The robots.txt guide includes a small named-agent example that shows why a wildcard policy can produce the wrong outcome.

Compare delivery with rendering

Start with what the server sent. Check the final URL, status, headers, content type, canonical, directives, main text, and ordinary links. Then compare it with the browser-rendered result and inspect failed resources or delayed network calls. A page can return 200 while serving a consent screen, error shell, empty application container, or the wrong canonical version.

Web components add a conditional branch to this investigation. If the page uses them, distinguish light DOM, a shadow root, slot assignment, and the composed browser result. An ordinary document.documentElement.outerHTML capture does not serialize shadow-tree contents. Google documents that its renderer can process light and shadow DOM content in its example; that is a Google-specific statement, not a general rule for other systems. (Google JavaScript SEO basics, DOM Standard: shadow trees)

Use the comparison playbook for the basic response-versus-rendered path. Its component-inspection branch is only needed when components, slots, or a discrepancy point to a shadow-tree issue.

Make indexability a separate decision

Indexability asks whether a processed page can be retained as the preferred version. noindex, X-Robots-Tag, redirects, canonical preferences, duplicate versions, and the content a crawler actually receives can all change that decision. A canonical is a consolidation signal; noindex is an instruction not to index. Blocking a URL before it is fetched may stop a crawler from seeing the page-level instruction, so access and index controls must be designed together.

Google says pages eligible to appear as supporting links in its AI features must be indexed and eligible to show a snippet. It does not require a special AI text file or additional structured data for that rule. That is a Google Search eligibility rule, not a universal condition for other answer products. (Google AI features and your website)

Record a conclusion that can be rechecked

For each material finding, retain the affected URL or template, tested stage, method and conditions, evidence, responsible owner, corrective action, and recheck condition. A useful conclusion says, for example, “the live response contains the canonical and page text, but the rendered product offer is missing after a failed API call,” rather than “the page is crawlable.”

The 29-point checklist supplies a working register and the full control set. It helps teams repeat the checks after a release, infrastructure change, crawler-policy revision, or a material observed failure.

Availability improves the chance that intended systems can inspect a page. It does not guarantee indexing, ranking, retrieval, AI-answer inclusion, citation, or a visit.