robots.txt · RFC 9309

robots.txt tester

Paste a robots.txt and find out what it actually says: which group applies to a given crawler, whether a URL is allowed or blocked, and exactly which rule decides it — with RFC 9309 longest-match precedence, wildcards, and percent-encoding handled the way real parsers handle them. Plus lint warnings for the classic mistakes. Nothing you paste leaves your browser.

How matching works

A crawler picks the group whose User-agent name matches it most specifically (longest matching name wins; * is only a fallback — a matching named group completely replaces it, never combines with it). Within the group, every rule is tried against the URL path: the rule with the longest pattern wins, and on a tie Allow beats Disallow. Rule order is irrelevant. * matches any run of characters, a trailing $ anchors the end of the URL, and patterns and paths are compared percent-encoded, so /café and /caf%C3%A9 are the same path — but %2F is not the same as a literal slash. If no rule matches, the URL is allowed: robots.txt is allow-by-default.

Remember what robots.txt is: a published request to polite crawlers, not access control. Anyone can read it — and a Disallow line pointing at a sensitive path is an advertisement of that path.