Audit Any Sitemap Instantly π
Inspect sitemap.xml files, sitemap indexes, discovered URLs, validation health, crawl signals and SEO-friendly details — fetched and parsed by our own backend, then status-checked URL by URL in batches, with anything not yet requested labelled as unchecked rather than assumed healthy.
Paste a full sitemap URL, or just a domain β if no sitemap URL is given we read
robots.txt and then check domain/sitemap.xml,
/sitemap_index.xml and other standard locations.
π§Sitemap Overview
Scan a sitemap to see how its files fit together.
Sitemap Structure
–π read successfully · β οΈ could not be read · every file opens in a new tab
Sitemap summary
URLs Found
–Discovered URLsLast Updated
–Response Status
–File Size
–UncompressedCompression
–π SEO Health Score
Rule-based, from real signalsπ Sitemap Files Detected
Real counts from the parsed XMLπ‘οΈ SEO & Crawl Details
Real XML & robots.txt signalsπ URL Coverage & Status
βπ‘ SEO Recommendations
π Performance Snapshot
Real, measuredURL breakdown
π§ͺ Status Breakdown
Checked URLs onlyClick any card to filter the URL list by that status.
ποΈ URL Types
From the sitemap itselfπ Top Sections
First path segmentπ§± URL Depth
Path segments per URLA Real Sitemap Audit, Not a Demo
Inspect fetches the sitemap URL from our own server with the same SSRF protections used across our tools (private/loopback/link-local/metadata addresses are always rejected), transparently detects and decompresses GZIP, and parses the XML as either a <urlset> or a <sitemapindex> (following one level of child sitemaps). It checks robots.txt for a sitemap reference and Disallow rules, then status-checks the URLs themselves with real HTTP requests — an instant first sample, then batches of 50 in the background — recording status codes, noindex directives and robots blocking. If you enter a bare domain instead of a sitemap URL, robots.txt and the standard sitemap locations are checked and the report says which one it used. Every count is either a real measurement or clearly labeled as not-yet-checked; nothing is randomly generated.
example.com while www.example.com also answers
is advertising one of two duplicate addresses, and the audit above cannot see that
— it only reads what the file says. The
Canonical Host Checker
tests all four variants and tells you which one your sitemap should be using.
Sitemap Checker FAQ
Does this check every URL in my sitemap?
It works towards it. The first scan validates a small evenly-spaced sample instantly, then the page keeps checking in batches of 50 β up to 500 URLs automatically β with the coverage bar filling in live. If the sitemap holds more than that, Check more URLs continues from where it stopped. Anything not yet requested stays grey and is labelled "Not checked" rather than assumed healthy.
Do I have to paste the full sitemap URL?
No. Enter a domain or any page URL and we read that site's robots.txt for a Sitemap: line, then try the standard locations β /sitemap.xml, /sitemap_index.xml, /wp-sitemap.xml and a few more. The report shows exactly which location was used.
How deep does it crawl a sitemap index?
One level. A sitemap index's child sitemaps are fetched and parsed for real; if a child is itself an index, its referenced count is shown but not recursively crawled.
Can I check internal or private URLs?
No. Requests that resolve to localhost, private IP ranges, link-local addresses or cloud metadata endpoints are rejected before any fetch happens.
What does "not exact" mean next to the URL count?
It means this scan hit its configured URL or child-sitemap limit while parsing, so the real total may be higher than what's shown.
Is my scanned data stored?
No. Each scan runs live for your request only. The exported report is generated and downloaded entirely in your browser.