Robots.txt and Sitemap Errors That Block Indexing
Sep 18, 2026 · 7 min read

Two plain text files decide whether Google can reach your site at all. Most companies have never opened either one. A plugin or a developer generated them at launch, and nobody has touched them since. Meanwhile a single line in robots.txt can keep an entire product category out of search results, and no notification will ever tell you.
Below are the mistakes I find in these files most often, and how to check for each one.
What each file actually does
robots.txt controls crawling. It tells bots where they may not go. It does not control what appears in search results. A blocked URL can still show up in Google — with no description, just a note that no information is available.
sitemap.xml is a list of URLs you consider important. It is a suggestion, not an instruction. A URL in the sitemap is not guaranteed to be indexed, and a URL outside the sitemap can be indexed anyway. Sitemaps help mostly on larger sites and on new pages that few links point to.
The practical consequence: robots.txt can do a lot of damage fast, a bad sitemap mostly wastes crawl budget quietly.
Robots.txt mistakes
Disallow: / left over from staging
The most expensive mistake I see in this file. The site was built on a staging domain where blocking everything was correct. At launch, the file was copied over along with the pages. The site looks fine, it just gradually disappears from results.
The check takes seconds: open yourdomain.com/robots.txt and look for a line reading Disallow: / with nothing after the slash.
Blocked CSS and JavaScript
Old advice said to block system directories. In practice that means Google downloads the page but cannot render it — no layout, no content pulled in by script. On WordPress this usually looks like Disallow: /wp-includes/ or blocking /wp-content/. Leave scripts and stylesheets accessible.
noindex inside robots.txt
A Noindex: directive in robots.txt does nothing. Google does not support it. If you need a page out of the index, that page must be crawlable and carry a meta robots tag with the value noindex, or an X-Robots-Tag header.
Disallow and noindex at the same time
This is the most common logical error. The page has noindex in its HTML and is also blocked in robots.txt. The crawler never visits the page, never sees the tag, and the URL can stay in the index indefinitely. If you want something gone from results, remove the robots.txt block first.
The same applies to canonicals. Nobody reads a blocked URL, so the canonical on it never counts — and the duplication you were trying to solve stays exactly where it was. More on that in canonical tags and duplicate content.
Wildcards that are too broad
A line like Disallow: /*? is meant to block ecommerce filters. What it actually blocks is every URL with a parameter — pagination, on-site search, any version of a page carrying UTM tags. If your pagination runs on parameters, you have just cut off the path to every product from page two onwards. In that case it is safer to handle filters with canonicals and internal links than with a blanket ban.
Rules do not merge
If the file has a group for User-agent: * and a separate group for User-agent: Googlebot, Googlebot follows only its own group. It ignores everything under the wildcard. I have seen sites where the Googlebot group contained one leftover line from an old test, and every other rule that was supposed to apply was simply never used.
The file is in the wrong place
robots.txt applies to one host, protocol and port only. https://yourdomain.com/robots.txt does not govern https://blog.yourdomain.com/. Every subdomain needs its own file in its own root. A copy in a subdirectory has no effect at all.
Server error responses
If robots.txt returns a 404, Google treats it as "nothing is disallowed" and crawls normally. If it returns a 5xx error, Google plays it safe and temporarily throttles crawling. Regular outages on this one file show up as unexplained slowdowns in indexing. Check that it returns status 200 with content type text/plain.
Crawl-delay
Google ignores the Crawl-delay directive. If server load is a real problem, fix it with response times or the crawl rate setting in Search Console, not with this line.
Sitemap mistakes
URLs that should not be indexed
Typical output from an automatic generator: pages with noindex, redirects, URLs whose canonical points somewhere else, tag archives, image attachment pages, the cart page. A sitemap should contain only final URLs returning 200 that you want in search results. Everything else is a contradictory signal.
404s and redirects
Deleted products stay in the sitemap for months. In Search Console this shows up as a growing count of discovered but not indexed URLs. The sitemap has to regenerate when content changes, not once at launch.
Wrong URL format
A sitemap needs absolute URLs in exactly the form the site runs on. Common mismatches: http:// instead of https://, the www version against the non-www one, a missing or extra trailing slash, relative paths. A sitemap also cannot list URLs from another domain.
lastmod that lies
If the generator writes today's date into lastmod for every URL on every run, the value stops meaning anything and gets ignored. lastmod should reflect a real content change. Google does not use priority or changefreq at all — there is nothing to tune there.
Limits exceeded
One sitemap file may contain at most 50,000 URLs and be at most 50 MB uncompressed. A larger site needs a sitemap index pointing at several smaller files. An ecommerce site with tens of thousands of products is better off splitting sitemaps by content type — products, categories, articles. In Search Console you then see which group has the problem.
Two sitemaps running at once
When switching SEO plugins, the old sitemap often keeps serving alongside the new one. Google gets two lists, one of them stale. Let the old one return 404 and remove it from the Search Console list.
The sitemap blocked in robots.txt
This happens with broad parameter rules or a rule on /sitemap. The sitemap has to be reachable. And the reference to it belongs in robots.txt as Sitemap: https://yourdomain.com/sitemap.xml — the one directive there that works regardless of user agent.
HTML returned instead of XML
If the server answers the sitemap URL with an error page, a redirect to the homepage, or a file with a BOM at the start, parsing fails. In Search Console it reads as "couldn't fetch sitemap".
A ten-minute check
- Open
yourdomain.com/robots.txtin a browser. Read every line. If you do not know what a line is for, that line needs checking. - In Search Console, open the robots.txt report under Settings. It shows when the file was last fetched, with what status, and whether Google found errors in it.
- Under Sitemaps, check the last read date and the number of discovered URLs. The gap between URLs in the sitemap and indexed pages is your to-do list.
- Take five important URLs — homepage, main category, top product, best-read article — and run each through URL Inspection. For each one check whether crawling is allowed, whether indexing is allowed, and which URL is set as canonical.
- Open the sitemap and pick ten random addresses. They all have to return 200 and carry no
noindex.
If the first and second both check out, the technical gate is open.
What this still does not solve
A correct robots.txt and a clean sitemap are not a reason to index anything. They are only the condition that makes indexing possible. If no internal link points to a page and the page has no content of its own, Google will happily leave it unindexed, sitemap or not. So after fixing these two files, the work moves to internal linking and to the content on individual pages.
Now, concretely, what I do and what I do not. I crawl the public pages of a site and report findings at page level — duplicate title tags, thin descriptions, missing alt text, broken structured data, misconfigured canonicals, slow pages. For the important findings I write the finished fix text, you approve it, and it gets deployed. I will not rewrite the contents of your robots.txt or restructure your sitemap — that is a change to server or plugin configuration and belongs to the person who maintains the site. The list above is exactly what you use to check it yourself or hand to a developer.
A reasonable start is a page-level audit, and after the fixes go live, measure the impact 28 days before against 28 days after. With technical blocking, though, one rule holds: as long as an unnecessary Disallow sits in robots.txt, the other fixes do not matter, because nobody can read them.
This is written by a tool you can buy
The article was proposed and written by Seonal — the same one that finds the errors on your site, fixes them and measures the result. The audit is free.