Robots.txt checker
Read a site's robots.txt and see which crawlers may visit which paths.
- Rules grouped by user agent
- Sitemap references
- Parse problems flagged
Give it something to look up — Glee will do the digging.
How the robots.txt checker works
robots.txt is the file at the root of a site that tells crawlers what not to fetch. One wrong line — a stray Disallow: / left over from a staging site — can remove a whole site from search results, and it is one of the first things to check when traffic drops.
This checker fetches /robots.txt, parses it into groups by user agent, and lists the allow and disallow rules, crawl delays and sitemap references it declares, with any lines a crawler would ignore.
What each field means
- Rules
- Allow and Disallow lines per user agent.
- Sitemaps
- Sitemap files the site declares.
Where this check stops
- A robots rule is not an indexing guarantee: a blocked page can still be indexed from links, and crawlers may ignore the file.
The same check, through the API
Same calculation, same answer, with a key. 5 credits per check ($0.50 per 1,000). A free account includes 1,000 credits a month.
POST /v1/web {"url":"acme.com"}
Host: gleanzy.com
Authorization: Bearer $GLEANZY_KEYFrequently asked
Does Disallow remove a page from search?
No — it stops crawling. Use a noindex directive on the page to keep it out.