Free Robots.txt Checker
Enter your website. We read your robots.txt, flag mistakes line by line, check the sitemaps it lists and show whether Google, Bing and AI crawlers may fetch any path you choose. No sign-up.
- 01Enter your website
- 02We read robots.txt
- 03See what each crawler may do
Who may fetch /
Check live AI crawler access →robots.txt is a request that well-behaved crawlers follow. It does not stop a page from appearing in search if other sites link to it; use noindex for that.
Checks
Sitemaps listed
theStacc’s site audit tests how search engines and AI crawlers actually reach every page of your site, not only what robots.txt asks.
Sign up for free →What is robots.txt?
robots.txt is a plain-text file at the root of a website, at yourdomain.com/robots.txt, that tells crawlers which paths they may fetch. Each group starts with one or more User-agent lines and lists Allow and Disallow rules for them. It can also point crawlers to your sitemaps. This checker reads the real file and applies the same matching rules crawlers use.
How it works
- Enter your website and, if you like, a path to test, such as /blog/ or /checkout/.
- We read robots.txt from your server, check it line by line and request each sitemap it lists.
- You see each crawler's answer for that path, Allowed or Blocked, with the rule and group that decided it, plus the file itself with problem lines highlighted.
Common robots.txt mistakes
- Disallow: / in the * group. Asks every crawler without its own group to stay away from the whole site. Often left over from a staging site.
- Rules before any User-agent line. They belong to no group, so no crawler applies them.
- Misspelled fields. "Dissallow" or "User agent" are ignored.
- Blocking pages you want removed from search. Use noindex instead; a blocked page cannot show crawlers its noindex tag.
- Blocking AI search crawlers by accident. A rule meant for training crawlers such as GPTBot can also catch OAI-SearchBot if it is written for *.
robots.txt vs noindex vs llms.txt
| robots.txt | noindex | llms.txt | |
|---|---|---|---|
| Controls | Which paths crawlers fetch | Whether a page appears in search | Nothing; it is a guide to your pages |
| Where | /robots.txt | Meta tag or HTTP header on the page | /llms.txt |
| Status | Internet standard (RFC 9309) | Supported by major search engines | Proposal (llmstxt.org) |
Frequently Asked Questions
It fetches /robots.txt from the site you enter, parses it the way crawlers do under the robots.txt standard (RFC 9309), and shows three things: mistakes in the file with their line numbers, whether each sitemap listed in it loads, and whether Googlebot, Bingbot and the main AI crawlers may fetch the path you choose, with the exact rule that decided it.
A crawler uses the group that names its own user agent; only if there is none does it use the * group. Within that group, the rule with the longest matching path wins, and when an Allow and a Disallow match equally, Allow wins. The checker applies the same logic.
Googlebot and Bingbot for search; OAI-SearchBot, Claude-SearchBot and PerplexityBot for AI search; ChatGPT-User, Claude-User and Perplexity-User for pages a user asks an assistant to open; and GPTBot, ClaudeBot, Google-Extended, Applebot-Extended and CCBot for AI training.
No. robots.txt asks crawlers not to fetch a page, but Google can still list a blocked URL if other pages link to it, usually without a description. To keep a page out of search results, let it be crawled and add a noindex meta tag or X-Robots-Tag header.
That is allowed. A robots.txt that returns 404 means crawlers may fetch everything. A robots.txt that returns a server error is different: Google treats it as 'crawl nothing' until it can read the file again.
No. Google ignores Crawl-delay (Bing respects it), and Google stopped supporting Noindex rules in robots.txt in 2019. The checker points out both.
No. We fetch the file when you click Check and keep nothing. The tool is free with no sign-up; each visitor can run a limited number of checks in a short period.
Last updated 2026-09-28.
Explore theStacc software
Checking robots.txt is free. Content Indexing keeps your pages discoverable by search and AI engines.
Sign up for free →