Skip to main content
← All guides

Robots.txt and sitemaps

Both files say something to a crawler about your URLs, and neither does what people usually assume. One withholds permission to fetch. The other offers a list to consider.

Permission is not the same as exclusion

robots.txt withholds permission to fetch a URL, and nothing else. Google states plainly that it is not a mechanism for keeping a page out of Search. The two ideas get collapsed constantly, and the collapse produces a specific, visible failure.

A disallowed URL that other sites link to can still be indexed. Google knows the address exists because someone linked to it; it just cannot see what is on it. So it may list the URL with no description at all — the worst of both worlds, since the page is public in the results while you have no control over how it appears.

Reading and writing the file

The file lives at the root of a host and applies only to that host and protocol. Rules are grouped by user-agent, and a crawler obeys the single most specific group that matches it — not every group it could match. Within a group, Google resolves conflicts by the most specific rule, not by order.

robots.txt
User-agent: *
Disallow: /admin/
Disallow: /*.pdf$
Allow: /admin/public/

User-agent: Googlebot
Disallow: /drafts/

Sitemap: https://example.com/sitemap.xml

Two details in that sample carry most of the confusion. The Allow line is more specific than the Disallow above it, so /admin/public/ stays fetchable. And once a Googlebot group exists, Googlebot obeys that group alone — the wildcard group no longer applies to it, so /admin/ is fetchable by Googlebot here. That is almost never what the author intended.

  • Blocking a path does not remove already-indexed URLs; it just stops future fetches.
  • The Sitemap directive is independent of any user-agent group and can appear anywhere in the file.
  • A 5xx on robots.txt itself can cause a crawler to back off the whole host, so serve it from something reliable.

Where a robots rule does damage nowhere near itself

The most expensive robots.txt mistakes are not the pages you meant to block. They are the resources a page needs in order to render. Engines render pages before indexing them, and rendering pulls scripts, styles and API responses.

Block a scripts path or an API route and the crawler still fetches the HTML perfectly well. It then renders a nearly empty page and indexes a nearly empty page. No status anywhere will say the page was blocked, because the page was not blocked — only the things it needed.

A sitemap is a suggestion, not a queue

Sitemaps help discovery where linking is weakest: large sites, brand-new sites with few inbound links, and sites heavy with media that has no natural anchor text. Google says outright that listing a URL guarantees neither crawling nor indexing.

Google's own guidance is that a smaller, well-linked site may not need one at all, because the crawler can find everything by following links. That is worth taking seriously before building sitemap infrastructure to solve a problem that internal linking would solve better.

LimitValue
URLs per sitemap file50,000
Uncompressed size per file50 MB
Beyond either limitSplit and reference the parts from a sitemap index

Which tags are read, ignored, or distrusted

This is where most sitemap effort is wasted. Two of the four common tags do nothing at all, and a third can actively work against you.

TagWhat Google does with it
locRead — the URL itself
lastmodUsed only when consistently and verifiably accurate
changefreqIgnored
priorityIgnored

The lastmod condition is the one that bites. If your build stamps every URL with the deploy timestamp, every value is simultaneously wrong and easy to disprove — the page did not change, and the engine can tell. Once it stops trusting the field it stops using it, which leaves you worse off than sending no lastmod at all. Emit it only for real changes to main content, and omit it rather than fake it.

sitemap.xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/guides/indexing/</loc>
    <lastmod>2026-06-14</lastmod>
  </url>
</urlset>

Choosing the right instrument

Almost every question in this area resolves once you separate three jobs that feel similar and are not.

GoalInstrumentWhy
Stop wasting crawl effort on junk URLsrobots.txtPrevents the fetch entirely
Keep a page out of the resultsnoindex, or authenticationRequires the fetch so the directive can be read
Help pages be discoveredInternal links, then a sitemapLinks carry context; sitemap entries do not

Questions and answers

Can I combine disallow and noindex?

No, and that pairing is the classic mistake. If robots.txt blocks the URL, the crawler never fetches the page and never sees the noindex. Allow the crawl and serve the noindex, or put the page behind authentication.

Do I need a sitemap at all?

Not always. Google's guidance is that a smaller, well-linked site can be discovered by following links. Sitemaps earn their place on large sites, brand-new sites with few inbound links, and rich media.

Why is a page I blocked still showing in Google?

Because robots.txt stopped the fetch, not the indexing. Other sites link to the URL, so Google knows it exists but cannot see the content — hence a listing with no description. Allow the crawl and serve a noindex to remove it properly.

Should I set priority and changefreq?

No. Google ignores both. The effort is better spent making lastmod accurate, or on internal linking.

What happens if robots.txt returns an error?

A 5xx can lead a crawler to back off the host rather than assume everything is allowed. Serve robots.txt from infrastructure at least as reliable as the site itself.