Home / Google News / Google SEO / Google Sitemaps Robots.txt Validator Does Not Properly Validate

Google Sitemaps Robots.txt Validator Does Not Properly Validate

Mar 23, 2006 - 8:00 am 1 — by Barry Schwartz

Filed Under Google Search Engine Optimization

Shawn Hogan of DigitalPoint wrote a blog entry named Google Not Interpreting robots.txt Consistently. He describes how he noticed that some of his pages were being crawled by GoogleBot, even though his robots.txt file specifically was blocking it. So he emailed Google, and they actually replied with the following message;

hile we normally don't review individual sites, we did examine your robots.txt file. Please be advised that it appears your Googlebot entry in your robots.txt file is overriding your generic User-Agent listing. We suggest you alter your robots.txt file by duplicating the forbidden paths under your Googlebot entry: User-agent: * Disallow: /tools/suggestion/? Disallow: /search.php Disallow: /go.php Disallow: /~shawn/scripts/ Disallow: /ads/
User-agent: Googlebot Disallow: /~shawn/ebay_ Disallow: /tools/suggestion/? Disallow: /search.php Disallow: /go.php Disallow: /~shawn/scripts/ Disallow: /ads/
Once you've altered your robots.txt file, Google will find it automatically after we next crawl your site.

Fine, so Shawn can easily do that. It is not a major deal, a bug Google knows about in its robots.txt protocol. But what Shawn points out is that the Google Sitemaps robots.txt validator shows that his previous robots.txt file;

User-agent: *
Disallow: /tools/suggestion/?
Disallow: /search.php
Disallow: /go.php
Disallow: /~shawn/scripts/
Disallow: /ads/
User-agent: Googlebot
Disallow: /~shawn/ebay_

was actually validated that it would not crawl the /ads/ directory. The two are not consistent, and should be, obviously.

Forum discussion at DigitalPoint Forums.

Previous Story: Ask.com Search Quality & Search Index Improving?

Next Story: Most, If Not All, Supplemental Google Index Issues Resolved

The content at the Search Engine Roundtable are the sole opinion of the authors and in no way reflect views of RustyBrick ®, Inc
Copyright © 1994-2025 RustyBrick ®, Inc. Web Development All Rights Reserved.
This work by Search Engine Roundtable is licensed under a Creative Commons Attribution 3.0 United States License. Creative Commons License and YouTube videos under YouTube's ToS.

Google Sitemaps Robots.txt Validator Does Not Properly Validate

Barry Schwartz / Executive Editor

Popular Categories

The Pulse of the search community

Search Video Recaps

Most Recent Articles

Daily Search Forum Recap: April 15, 2025

Financial Times Interviews Head Of Google Search, Elizabeth Reid

Google AI Overviews Linking Over & Over Again To Itself

Google Again Says Structured Data Does Not Make Your Site Rank Better

Google Bug Product Snippet Image Carousel

Google Tests AI Mode Shortcut In Search Bar