Resources / Crawlability, Rendering & Indexability / Tutorial

Build and Validate an XML Sitemap

A practical XML sitemap tutorial for choosing URLs, generating valid files, submitting them, and detecting stale or contradictory entries.

What a sitemap contributes

An XML sitemap gives crawlers a maintained list of URLs you want them to discover. It is especially useful for a new site, a large site, pages with few internal links, or a site with frequent additions. Submission is a discovery signal; it does not guarantee crawling, indexing, ranking, or inclusion in an AI answer.

Use the sitemap as a clean expression of the site’s preferred public URLs. Do not treat it as a dump of every address your platform can generate.

Start with one valid file

For a small site, create /sitemap.xml at the site root. Every <loc> value should be a complete absolute URL using the intended protocol and host.

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://www.example.com/</loc>
    <lastmod>2026-09-18</lastmod>
  </url>
  <url>
    <loc>https://www.example.com/resources/</loc>
    <lastmod>2026-09-16</lastmod>
  </url>
</urlset>

<loc> is required. <lastmod> is optional, but when present it should reflect a meaningful change to the page rather than the time the sitemap was regenerated. Google says it ignores <priority> and <changefreq>, so omitting them keeps the file simpler.

Recognize one common invalid entry

This URL contains an unescaped ampersand and will fail XML parsing:

<loc>https://www.example.com/catalog?color=blue&size=medium</loc>

Encode the XML-reserved character as &amp; while keeping the actual URL query parameter unchanged:

<loc>https://www.example.com/catalog?color=blue&amp;size=medium</loc>

The protocol requires XML escaping in the file; the crawler reads the intended URL after XML parsing. (Sitemaps.org protocol)

Choose URLs deliberately

Include a URL when all of these statements are true:

Exclude redirects, error pages, internal search results, parameter duplicates, staging URLs, pages marked noindex, and alternate versions that canonicalize elsewhere. A sitemap that contradicts canonical tags or indexing controls creates ambiguity rather than clarity.

Split large sitemaps with an index

A single sitemap is limited to 50,000 URLs or 50 MB uncompressed. Split a larger set into stable groups and list those files in a sitemap index.

<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://www.example.com/sitemaps/pages.xml</loc>
    <lastmod>2026-09-18</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://www.example.com/sitemaps/products.xml</loc>
    <lastmod>2026-09-18</lastmod>
  </sitemap>
</sitemapindex>

Group files by a dimension you can maintain and diagnose, such as content type, locale, or publication period. Stable grouping makes errors easier to isolate in search-engine reports.

Generate and publish it

  1. Export canonical, index-eligible URLs from the publishing system or route manifest.
  2. Normalize each URL to the public HTTPS hostname and preferred path form.
  3. Remove duplicates and any URL that redirects, errors, requires authentication, or carries noindex.
  4. Add <lastmod> only when the source system can supply a reliable substantive modification date.
  5. Escape XML-reserved characters in URLs, encode the file as UTF-8, and generate either one sitemap or a sitemap index.
  6. Publish the file on the live site. A root-level location gives it the broadest default scope.
  7. Add its absolute URL to robots.txt:
Sitemap: https://www.example.com/sitemap.xml
  1. Submit the sitemap or sitemap index through each search engine’s supported webmaster console. Submission complements the public file and robots.txt reference; it does not replace them.

Validate the live file

Check the response before checking the XML:

curl -I https://www.example.com/sitemap.xml
curl -L https://www.example.com/sitemap.xml | xmllint --noout -

Confirm that the URL returns a successful response, the body is XML rather than an HTML error or login page, and the file is available without cookies. If xmllint is unavailable, use another XML parser rather than relying only on a browser’s formatted display.

Then sample URLs from every sitemap group. Confirm status, canonical tag, meta robots directive, hostname, and internal discoverability. Compare the complete sitemap inventory with the canonical URL inventory from your CMS or crawl; sampling alone will miss systemic omissions.

Finally, review the sitemap report in each search console after processing. Separate syntax or fetch errors from index coverage decisions. A submitted URL can be discovered successfully and still remain unindexed for another reason.

Maintenance checks

Continue the technical review

Use the robots.txt guide and templates to keep crawl policy consistent with the sitemap. Use the raw-versus-rendered HTML playbook to inspect what a crawler can process at the listed URLs.

Official references