Topics / Crawlability, Rendering & Indexability
Crawlability, Rendering & Indexability
Important pages can fail before anyone evaluates their content. A system may not discover a URL, may be unable or unwilling to request it, may receive the wrong response, may render an incomplete page, or may decide not to retain the resulting URL. Those are different failures, so they need different evidence and different fixes.
Start with the resource that matches the work in front of you. The technical review checklist records a complete audit. The robots.txt guide turns a crawler-policy decision into rules. The XML sitemap tutorial builds a maintained discovery file. The original-versus-rendered HTML playbook investigates a page whose browser result differs from its delivered response.
This subject matters to AI SEO because an answer system can only use information it can obtain through its own documented access and processing path. Passing a Google test does not establish how another crawler, search product, or user-requested fetcher behaves.
Start with the demonstrated failure
When a material URL has an evidenced access, delivery, rendering, or indexability problem, investigate and recheck that stage before adding an alternate format or expanding discovery signals. If no failure is demonstrated, choose the smallest representative test from the symptom table. The first task is complete when the recorded evidence shows the intended stage working under the relevant conditions; then consider dependent improvements.
On this page
- Trace the page through six stages
- Choose the first test from the symptom
- Separate crawler purpose from crawler access
- Compare delivery with rendering
- Make indexability a separate decision
- Record a conclusion that can be rechecked
Trace the page through six stages
| Stage | Question | Useful evidence | It does not establish |
|---|---|---|---|
| Discovery | Can a system learn the URL exists? | Crawlable links, sitemaps, feeds, external links | That it requested the URL |
| Access | Is the intended agent permitted and technically able to request it? | robots.txt, authentication and edge policy, verified logs | That the response was usable |
| Delivery | Did the server return the intended page? | Final URL, status, headers, response body, timing | That browser-dependent content appeared |
| Rendering | Which content, links, and signals exist after processing? | Original HTML, rendered DOM, resource failures, inspection output | That the page will be indexed |
| Indexability | Is this version eligible to be retained as the preferred URL? | noindex, canonical signals, redirects, duplicate analysis | That it will be selected for a result |
| Retrieval and use | Was it selected for a search result, answer, citation, or visit? | Product reports, answer captures, result observation, referrals | Why it was selected |
Google documents crawling, rendering, and indexing as distinct stages for JavaScript pages. Its documentation is strong evidence about Google, including its own renderer, but it does not describe every search crawler or AI product. (Google JavaScript SEO basics)
Choose the first test from the symptom
| Symptom | First test | Next resource |
|---|---|---|
| A crawler should be allowed, but requests are denied or uncertain | Check named-agent policy, identity verification, edge rules, and a verified request | robots.txt for Search and AI crawlers |
| Important URLs are missing from discovery paths | Check internal links, canonical URLs, and sitemap inclusion | Build and validate an XML sitemap |
| A browser displays content that the initial response lacks | Save and compare the response, rendered DOM, links, metadata, and JSON-LD | Compare original and rendered HTML |
| An accessible page is excluded or consolidated elsewhere | Inspect directives, canonicalization, redirects, duplicates, and the relevant platform report | Crawlability checklist |
| A citation, mention, or referral is absent | Capture the outcome first, then test the earlier stage supported by evidence | Where to start with an AI SEO problem |
Do not use a missing citation as proof of a crawl block, or a successful 200 as proof of indexability. The symptom determines the first layer to inspect.
Separate crawler purpose from crawler access
One provider may operate different agents for conventional search, AI search, training collection, and a fetch triggered by a user. The policy decision is therefore specific: which public material may this named agent request, for which documented purpose?
robots.txt gives instructions to cooperating crawlers. It is not authentication, a way to conceal sensitive data, or a reliable removal mechanism for an already known web-page URL. Google specifically advises using noindex or access control when the aim is to keep a web page out of Google Search. (Google’s robots.txt guide)
Before editing a rule, record the public paths in scope, exact user agent, documented purpose, the provider’s identity-verification method, policy owner, and recheck trigger. The robots.txt guide includes a small named-agent example that shows why a wildcard policy can produce the wrong outcome.
Compare delivery with rendering
Start with what the server sent. Check the final URL, status, headers, content type, canonical, directives, main text, and ordinary links. Then compare it with the browser-rendered result and inspect failed resources or delayed network calls. A page can return 200 while serving a consent screen, error shell, empty application container, or the wrong canonical version.
Web components add a conditional branch to this investigation. If the page uses them, distinguish light DOM, a shadow root, slot assignment, and the composed browser result. An ordinary document.documentElement.outerHTML capture does not serialize shadow-tree contents. Google documents that its renderer can process light and shadow DOM content in its example; that is a Google-specific statement, not a general rule for other systems. (Google JavaScript SEO basics, DOM Standard: shadow trees)
Use the comparison playbook for the basic response-versus-rendered path. Its component-inspection branch is only needed when components, slots, or a discrepancy point to a shadow-tree issue.
Make indexability a separate decision
Indexability asks whether a processed page can be retained as the preferred version. noindex, X-Robots-Tag, redirects, canonical preferences, duplicate versions, and the content a crawler actually receives can all change that decision. A canonical is a consolidation signal; noindex is an instruction not to index. Blocking a URL before it is fetched may stop a crawler from seeing the page-level instruction, so access and index controls must be designed together.
Google says pages eligible to appear as supporting links in its AI features must be indexed and eligible to show a snippet. It does not require a special AI text file or additional structured data for that rule. That is a Google Search eligibility rule, not a universal condition for other answer products. (Google AI features and your website)
Record a conclusion that can be rechecked
For each material finding, retain the affected URL or template, tested stage, method and conditions, evidence, responsible owner, corrective action, and recheck condition. A useful conclusion says, for example, “the live response contains the canonical and page text, but the rendered product offer is missing after a failed API call,” rather than “the page is crawlable.”
The 29-point checklist supplies a working register and the full control set. It helps teams repeat the checks after a release, infrastructure change, crawler-policy revision, or a material observed failure.
Availability improves the chance that intended systems can inspect a page. It does not guarantee indexing, ranking, retrieval, AI-answer inclusion, citation, or a visit.