Sitemap Validator

Check an XML sitemap for errors and warnings against the sitemap protocol.

How to use Sitemap Validator

  1. 1Paste your raw XML sitemap code or paste an active public sitemap URL.
  2. 2Click 'Validate Sitemap' to execute structural syntax and schema verification.
  3. 3Inspect the validation dashboard for errors, warnings, and protocol compliance alerts.
  4. 4Review specific line-by-line syntax issues, such as unescaped characters or invalid date strings.
  5. 5Correct detected issues and re-validate before submitting to Google Search Console.

Key Features & Highlights

  • •Strict Sitemaps.org 0.9 Schema Compliance: Validates root <urlset> or <sitemapindex> structures, namespaces, and required child tags.
  • •ISO 8601 Lastmod Date Verification: Identifies invalid or malformed timestamp strings (YYYY-MM-DD or full UTC ISO timestamps).
  • •URL Format & Entity Escaping Checks: Detects unescaped ampersands (&), quotes, and non-ASCII characters that crash XML parsers.
  • •Quota & Limit Enforcement: Verifies your file does not exceed Google's 50,000 URLs or 50MB file size limits.
  • •100% In-Browser Privacy: Sensitive development staging sitemaps are verified locally without third-party data collection.

Understanding Sitemap Validator

A single syntax error, unescaped ampersand, or malformed XML tag can cause Googlebot and Bingbot to reject an entire sitemap file, stalling indexation across your entire website. Our free online Sitemap Validator parses and tests your XML code against official W3C and Sitemaps.org 0.9 schemas, highlighting critical syntax defects, broken loc tags, and invalid lastmod timestamps with precise diagnostics.

Frequently Asked Questions

What are the most common XML sitemap validation errors?

The most frequent sitemap errors are unescaped ampersands in URLs (using '&' instead of the required XML entity '&amp;'), invalid date formats in the <lastmod> tag (which must follow W3C ISO 8601), missing XML namespace declarations (<urlset xmlns='...'>), and non-UTF-8 character encoding issues.

Why does Google Search Console show 'Sitemap couldn't be read' or 'Couldn't fetch'?

'Couldn't fetch' or 'Sitemap couldn't be read' in Google Search Console typically occurs because the sitemap URL is blocked by robots.txt, returns a non-200 HTTP status (such as a 301 redirect or 403 forbidden), times out due to server latency, or contains fatal XML syntax errors that crash Google's parser.

What date format is required for the <lastmod> tag?

The <lastmod> tag must follow the W3C Datetime format (a subset of ISO 8601). Valid formats include either a simple date string (YYYY-MM-DD, e.g. 2026-05-15) or a complete date-time string with time zone designator (e.g. 2026-05-15T09:45:00+00:00). Using formats like 'MM/DD/YYYY' or invalid timestamps will trigger crawler warnings.

Can an XML sitemap contain relative URLs?

No. All URLs inside the <loc> tag must be complete, absolute canonical URLs including the protocol (https://) and domain name (e.g. https://example.com/page-1). Relative paths such as '/page-1' are strictly invalid per sitemaps.org standards and will be ignored by search engine crawlers.

How do I escape special characters in XML sitemap URLs?

XML requires 5 predefined entities to be escaped: ampersand (&) must be written as '&amp;', single quote (') as '&apos;', double quote (") as '&quot;', greater than (>) as '&gt;', and less than (<) as '&lt;'. URLs containing query parameters like '?page=1&sort=desc' must always be escaped as '?page=1&amp;sort=desc'.

Related Free Tools

Explore complementary utilities to boost your workflow

View all free tools →