Resources / Crawlability, Rendering & Indexability / Playbook

Compare Original and Rendered HTML

A repeatable playbook for finding content, links, metadata, and structured data that disappear or change when a page depends on JavaScript.

What this comparison shows

The original HTML is the server response before page scripts run. The rendered document is the page state after a browser has executed scripts and changed the DOM. Comparing them reveals which important information depends on rendering and whether that process introduces omissions, contradictions, or failures.

One serialization does not describe every tree a browser uses. Light DOM consists of a component host’s ordinary child nodes. A shadow tree is attached to the host and can contain its own elements and slots; assigned light DOM can appear through those slots in the composed page. document.documentElement.outerHTML records the document tree, including light DOM and shadow hosts, but it does not include the contents of their shadow roots.

The two versions do not need to be byte-for-byte identical. Timestamps, generated identifiers, consent markup, and personalization can create harmless noise. The useful question is whether a crawler or agent can retrieve and understand the page’s essential content and signals.

Fictional page · same URL, two captures
1. Original HTTP response
<h1>Team plan</h1>
<div id="price"></div>
<script src="plan.js">
</script>
Price text absent. A client reading only this HTML cannot extract the price from it.
2. DOM after JavaScript
<h1>Team plan</h1>
<div id="price">
  USD 30 / month
</div>
Price text present. In this example, a script inserts it into the browser’s document.
Finding: the price depends on renderingPreserve both captures. Verify the intended consumer’s actual output. If it needs price in the initial response, make that fact available there and repeat the comparison.
These are simplified excerpts, not complete documents or a shadow-DOM example. Seeing a price in your browser does not prove that a crawler received or rendered it.

Start with the basic comparison

Use steps 1–3 and 5 before investigating component internals. Capture one clean HTTP response, capture the settled browser document without interaction, and compare the content and signals that matter. The shadow-DOM branch in step 4 is conditional: use it only when the page uses web components or the basic comparison points to component-supplied content.

1. Choose representative pages

Select at least one URL from every important template: homepage, category, article, product, location, and any page with interactive or lazy-loaded content. Add known problem cases and a simple server-rendered page as a control.

For each URL, write down what must survive rendering: title, main heading, primary copy, important links, canonical URL, robots directives, language annotations, and structured data. Note any custom elements that supply or project essential content, controls, or links. This turns a vague visual comparison into a test with explicit expectations.

2. Capture the HTTP response

Save headers and the final response body from a clean request:

mkdir -p rendering-check
curl -sS -L -D rendering-check/headers.txt \
  -o rendering-check/original.html \
  https://www.example.com/page/

Review the status and redirect chain in headers.txt. Open original.html and search for the expected text, links, canonical tag, robots meta tag, and JSON-LD. This file shows what a non-rendering client receives.

If the page varies by device, locale, cookie, or user agent, repeat the capture with those conditions recorded. Do not label one request as universally representative.

3. Capture the rendered document

Open the same public URL in a private browser window with the developer tools available. Wait for the normal initial load to settle without clicking, scrolling, or accepting optional personalization. Check the Console and Network panels for blocked scripts, failed API calls, timeouts, and challenge responses.

In the browser console, copy the current rendered DOM:

copy(document.documentElement.outerHTML)

Paste the result into rendering-check/rendered.html. This captures the current document tree, not the original server response or the complete composed page. If browser automation is already part of your test stack, use it to save the same post-load document across templates and releases.

4. Conditional branch: inspect light DOM, shadow roots, and slots

Search the rendered document and the browser’s Elements panel for custom elements. For each component that contains essential information or navigation, record:

Run this diagnostic in the browser console to inventory open shadow roots, including open roots nested inside other open roots:

const shadowInventory = [];

function inspectRoot(root, location = "document") {
  for (const element of root.querySelectorAll("*")) {
    if (!element.shadowRoot) continue;

    const host = `${element.localName}${element.id ? `#${element.id}` : ""}`;
    const shadowLocation = `${location} > ${host}::shadow-root`;

    shadowInventory.push({
      host,
      location: shadowLocation,
      mode: element.shadowRoot.mode,
      text: element.shadowRoot.textContent?.trim() || "",
      html: element.shadowRoot.innerHTML,
      slots: [...element.shadowRoot.querySelectorAll("slot")].map((slot) => ({
        name: slot.name || "(default)",
        assignedElements: slot.assignedElements({ flatten: true }).map(
          (assigned) => `${assigned.localName}${assigned.id ? `#${assigned.id}` : ""}`,
        ),
      })),
    });

    inspectRoot(element.shadowRoot, shadowLocation);
  }
}

inspectRoot(document);
copy(JSON.stringify(shadowInventory, null, 2));

Save the output as rendering-check/open-shadow-roots.json. The script cannot enter a closed shadow root because ordinary outside page script does not receive that root through element.shadowRoot. Record closed-root hosts visible in developer tools when available, but do not interpret closed as evidence that a search renderer either can or cannot process the component.

Light DOM remains under the host in the document even when a slot projects it elsewhere in the composed page. Check slot assignment with the inventory rather than counting the same node as separate content. If fallback content appears only when a slot has no assigned nodes, test both the intended assignment and the fallback state.

5. Compare the meaningful elements

Start with a line diff:

diff -u rendering-check/original.html rendering-check/rendered.html

Minified markup may produce an unreadable diff. Format both files with the same HTML formatter, or extract and compare the specific elements that matter. Review these groups in order:

  1. Primary content: page title, main heading, explanatory text, prices, availability, dates, and other facts.
  2. Discovery: ordinary <a href> links to important pages and assets.
  3. Indexing controls: canonical URL, meta robots, language and regional annotations.
  4. Structured data: JSON-LD nodes, identifiers, offers, and values that should match the visible page.
  5. Web components: essential content in light DOM, open shadow roots, assigned slots, and the composed result.
  6. Resources: scripts, styles, images, fonts, and API requests required to produce the result.

Classify each important element as present in both versions, added during rendering, changed during rendering, or missing after rendering. For component content, also record whether it originates in light DOM or a shadow root and whether a slot projects it into the composed page. Content added during rendering deserves extra testing because fetchers differ in their ability and willingness to execute JavaScript.

Worked comparison

Suppose a product page’s response includes the h1, price, canonical, and Product JSON-LD, and the rendered document adds only a nonessential stock-animation container. Record the main content, link, canonical, and JSON-LD as present in both; record the animation as added during rendering. No rendering defect follows from that difference.

If instead the response contains only <div id="product"></div> and the rendered document adds the h1, price, and offer after an API request, classify those facts as added during rendering. The next decision is to check the request, the renderer evidence for the intended consumer, and whether the source can supply the important facts before JavaScript runs.

6. Test dependency failures

Repeat the rendered test while looking for realistic failure modes:

Google documents separate crawling, rendering, and indexing stages and says it does not render JavaScript from blocked pages or files. Other search crawlers and AI fetchers may process JavaScript differently, so a successful desktop browser load proves only that the browser completed that test.

7. Compare with a search-engine rendering

For an indexable production URL, run Google Search Console’s live URL test. Its tested-page details can show the raw HTML, HTTP headers, loaded resources, JavaScript console output, and a rendered screenshot. Google says its renderer flattens light DOM and shadow DOM into rendered HTML, so check that output for the component’s essential content rather than assuming the local outerHTML capture is equivalent.

Compare the search-engine evidence with the light DOM, open-shadow-root inventory, composed browser display, and accessibility information available in your local test. The inspection result represents Google-InspectionTool, not every search crawler or AI agent.

Record the testing time and whether you reviewed a live test or an indexed version. Cached index evidence may predate the current deployment.

8. Record and prioritize the finding

Use one record per mismatch:

URL: https://www.example.com/page/
Template: Product detail
Expected element: Current price and Offer JSON-LD
Original HTML: Missing
Light DOM: Price placeholder only
Open shadow root: Current price present after product API request
Slot or composed result: Price visible; no slot involved
Google rendered HTML: Current price present
Failure observed: API returned 403 to the test client
Impact: Price and offer unavailable when rendering fails
Owner: Web platform
Next check: Retest after server-rendered price release

Prioritize a mismatch when it affects the main answer, internal discovery, indexability, canonicalization, or factual consistency. Treat decorative differences and ephemeral attributes as noise unless they reveal a wider failure.

9. Recheck after the fix

Retest the same URLs and conditions after deployment. Confirm both the intended correction and the absence of new differences in adjacent templates. Keep the before-and-after evidence with the issue so later changes can be compared against a known result.

Continue the technical review

Use the robots.txt guide to verify that the page and its dependencies are crawlable. Use the XML sitemap tutorial to confirm that preferred, index-eligible URLs are discoverable.

Official references