Permission is not the same as exclusion
robots.txt withholds permission to fetch a URL, and nothing else. Google states plainly that it is not a mechanism for keeping a page out of Search. The two ideas get collapsed constantly, and the collapse produces a specific, visible failure.
A disallowed URL that other sites link to can still be indexed. Google knows the address exists because someone linked to it; it just cannot see what is on it. So it may list the URL with no description at all — the worst of both worlds, since the page is public in the results while you have no control over how it appears.
Reading and writing the file
The file lives at the root of a host and applies only to that host and protocol. Rules are grouped by user-agent, and a crawler obeys the single most specific group that matches it — not every group it could match. Within a group, Google resolves conflicts by the most specific rule, not by order.
User-agent: *
Disallow: /admin/
Disallow: /*.pdf$
Allow: /admin/public/
User-agent: Googlebot
Disallow: /drafts/
Sitemap: https://example.com/sitemap.xml Two details in that sample carry most of the confusion. The Allow line is more specific than the Disallow above it, so /admin/public/ stays fetchable. And once a Googlebot group exists, Googlebot obeys that group alone — the wildcard group no longer applies to it, so /admin/ is fetchable by Googlebot here. That is almost never what the author intended.
- Blocking a path does not remove already-indexed URLs; it just stops future fetches.
- The Sitemap directive is independent of any user-agent group and can appear anywhere in the file.
- A 5xx on robots.txt itself can cause a crawler to back off the whole host, so serve it from something reliable.
Where a robots rule does damage nowhere near itself
The most expensive robots.txt mistakes are not the pages you meant to block. They are the resources a page needs in order to render. Engines render pages before indexing them, and rendering pulls scripts, styles and API responses.
Block a scripts path or an API route and the crawler still fetches the HTML perfectly well. It then renders a nearly empty page and indexes a nearly empty page. No status anywhere will say the page was blocked, because the page was not blocked — only the things it needed.
A sitemap is a suggestion, not a queue
Sitemaps help discovery where linking is weakest: large sites, brand-new sites with few inbound links, and sites heavy with media that has no natural anchor text. Google says outright that listing a URL guarantees neither crawling nor indexing.
Google's own guidance is that a smaller, well-linked site may not need one at all, because the crawler can find everything by following links. That is worth taking seriously before building sitemap infrastructure to solve a problem that internal linking would solve better.
| Limit | Value |
|---|---|
| URLs per sitemap file | 50,000 |
| Uncompressed size per file | 50 MB |
| Beyond either limit | Split and reference the parts from a sitemap index |
Which tags are read, ignored, or distrusted
This is where most sitemap effort is wasted. Two of the four common tags do nothing at all, and a third can actively work against you.
| Tag | What Google does with it |
|---|---|
| loc | Read — the URL itself |
| lastmod | Used only when consistently and verifiably accurate |
| changefreq | Ignored |
| priority | Ignored |
The lastmod condition is the one that bites. If your build stamps every URL with the deploy timestamp, every value is simultaneously wrong and easy to disprove — the page did not change, and the engine can tell. Once it stops trusting the field it stops using it, which leaves you worse off than sending no lastmod at all. Emit it only for real changes to main content, and omit it rather than fake it.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/guides/indexing/</loc>
<lastmod>2026-06-14</lastmod>
</url>
</urlset> Choosing the right instrument
Almost every question in this area resolves once you separate three jobs that feel similar and are not.
| Goal | Instrument | Why |
|---|---|---|
| Stop wasting crawl effort on junk URLs | robots.txt | Prevents the fetch entirely |
| Keep a page out of the results | noindex, or authentication | Requires the fetch so the directive can be read |
| Help pages be discovered | Internal links, then a sitemap | Links carry context; sitemap entries do not |