XML sitemap checker
Find a site's XML sitemap, check that it parses and see how many pages it lists.
- Location from robots.txt or the usual path
- Parse errors explained
- URL count, bounded
Give it something to look up — Glee will do the digging.
How the sitemap checker works
A sitemap lists the pages a site wants search engines to know about. A broken one is silently ignored, and a missing one leaves new pages to be found by links alone.
This checker looks for the sitemap where robots.txt declares it, and otherwise at the conventional /sitemap.xml, downloads it within a size limit, and reports whether it parses, whether it is an index of further sitemaps, and how many URLs it lists.
Common faults are easy to miss by eye: a sitemap that lists http:// addresses on an https:// site, one served with an HTML error page instead of XML, or an index pointing at files that no longer exist. Each is reported as a finding, so the fix is clear.
What each field means
- Status
- Found and valid, found with errors, or not found.
- URL count
- Pages listed, counted up to the download limit.
Where this check stops
- Downloads and parsing are bounded; finding a sitemap is not proof that the pages are indexed.
The same check, through the API
Same calculation, same answer, with a key. 5 credits per check ($0.50 per 1,000). A free account includes 1,000 credits a month.
POST /v1/web {"url":"acme.com"}
Host: gleanzy.com
Authorization: Bearer $GLEANZY_KEYFrequently asked
Do I need a sitemap?
Small, well-linked sites can do without; large or new sites benefit.
How many URLs can one sitemap hold?
Up to 50,000 URLs or 50 MB uncompressed; larger sites split them and list the parts in a sitemap index.