Resources / Crawlability, Rendering & Indexability / Tutorial
Build and Validate an XML Sitemap
A practical XML sitemap tutorial for choosing URLs, generating valid files, submitting them, and detecting stale or contradictory entries.
What a sitemap contributes
An XML sitemap gives crawlers a maintained list of URLs you want them to discover. It is especially useful for a new site, a large site, pages with few internal links, or a site with frequent additions. Submission is a discovery signal; it does not guarantee crawling, indexing, ranking, or inclusion in an AI answer.
Use the sitemap as a clean expression of the site’s preferred public URLs. Do not treat it as a dump of every address your platform can generate.
Start with one valid file
For a small site, create /sitemap.xml at the site root. Every <loc> value should be a complete absolute URL using the intended protocol and host.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/</loc>
<lastmod>2026-09-18</lastmod>
</url>
<url>
<loc>https://www.example.com/resources/</loc>
<lastmod>2026-09-16</lastmod>
</url>
</urlset>
<loc> is required. <lastmod> is optional, but when present it should reflect a meaningful change to the page rather than the time the sitemap was regenerated. Google says it ignores <priority> and <changefreq>, so omitting them keeps the file simpler.
Recognize one common invalid entry
This URL contains an unescaped ampersand and will fail XML parsing:
<loc>https://www.example.com/catalog?color=blue&size=medium</loc>
Encode the XML-reserved character as & while keeping the actual URL query parameter unchanged:
<loc>https://www.example.com/catalog?color=blue&size=medium</loc>
The protocol requires XML escaping in the file; the crawler reads the intended URL after XML parsing. (Sitemaps.org protocol)
Choose URLs deliberately
Include a URL when all of these statements are true:
- It is public and returns the intended content with a successful status.
- It is the preferred canonical version of the page.
- It is eligible for indexing under the site’s policy.
- Internal links and redirects support the same preferred URL.
- Its protocol, hostname, path, capitalization, and trailing-slash form are correct.
Exclude redirects, error pages, internal search results, parameter duplicates, staging URLs, pages marked noindex, and alternate versions that canonicalize elsewhere. A sitemap that contradicts canonical tags or indexing controls creates ambiguity rather than clarity.
Split large sitemaps with an index
A single sitemap is limited to 50,000 URLs or 50 MB uncompressed. Split a larger set into stable groups and list those files in a sitemap index.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/sitemaps/pages.xml</loc>
<lastmod>2026-09-18</lastmod>
</sitemap>
<sitemap>
<loc>https://www.example.com/sitemaps/products.xml</loc>
<lastmod>2026-09-18</lastmod>
</sitemap>
</sitemapindex>
Group files by a dimension you can maintain and diagnose, such as content type, locale, or publication period. Stable grouping makes errors easier to isolate in search-engine reports.
Generate and publish it
- Export canonical, index-eligible URLs from the publishing system or route manifest.
- Normalize each URL to the public HTTPS hostname and preferred path form.
- Remove duplicates and any URL that redirects, errors, requires authentication, or carries
noindex. - Add
<lastmod>only when the source system can supply a reliable substantive modification date. - Escape XML-reserved characters in URLs, encode the file as UTF-8, and generate either one sitemap or a sitemap index.
- Publish the file on the live site. A root-level location gives it the broadest default scope.
- Add its absolute URL to
robots.txt:
Sitemap: https://www.example.com/sitemap.xml
- Submit the sitemap or sitemap index through each search engine’s supported webmaster console. Submission complements the public file and
robots.txtreference; it does not replace them.
Validate the live file
Check the response before checking the XML:
curl -I https://www.example.com/sitemap.xml
curl -L https://www.example.com/sitemap.xml | xmllint --noout -
Confirm that the URL returns a successful response, the body is XML rather than an HTML error or login page, and the file is available without cookies. If xmllint is unavailable, use another XML parser rather than relying only on a browser’s formatted display.
Then sample URLs from every sitemap group. Confirm status, canonical tag, meta robots directive, hostname, and internal discoverability. Compare the complete sitemap inventory with the canonical URL inventory from your CMS or crawl; sampling alone will miss systemic omissions.
Finally, review the sitemap report in each search console after processing. Separate syntax or fetch errors from index coverage decisions. A submitted URL can be discovered successfully and still remain unindexed for another reason.
Maintenance checks
- Regenerate the sitemap when publishable URLs are added, removed, moved, or canonically changed.
- Remove deleted and redirected URLs instead of keeping them as a historical ledger.
- Update
<lastmod>only after a meaningful page change. - Monitor sitemap fetch failures and unexpected drops in submitted or discovered URLs.
- Recheck generation after migrations, hostname changes, locale launches, or CMS updates.
Continue the technical review
Use the robots.txt guide and templates to keep crawl policy consistent with the sitemap. Use the raw-versus-rendered HTML playbook to inspect what a crawler can process at the listed URLs.