Skip to main content
← All guides

Why pages don't get indexed

A page can stay out of Google's index for several different reasons: access or directives, duplicate-page or canonical handling, a crawl that has not happened yet, or a crawl that did not lead to indexing. Search Console reports the indexing status; the fix depends on which status you actually have.

Start with the status, not a diagnosis

The Page indexing report includes both genuine technical problems and expected exclusions. A robots.txt rule, noindex directive, authentication requirement or failing response can prevent indexing; duplicate and canonical statuses may be working exactly as intended; and a discovered URL may simply not have been crawled yet.

The most misleading state is “Crawled — currently not indexed”. It confirms that Google fetched the page but did not add it to the index at that time. Google explicitly says the page may or may not be indexed later and that there is no need to resubmit it just because of this status.

The statuses that are genuinely errors

These describe something on your side that stopped the page being stored. Google labels the source of each as Website when you can fix it yourself, and Google when the cause sits on their side. Correcting the underlying response is the whole fix.

StatusWhat it meansSource
Server error (5xx)The server returned a 500-level errorGoogle
Redirect errorA redirect chain looped or otherwise failedWebsite
URL blocked by robots.txtrobots.txt withheld permission to fetchWebsite
URL marked ‘noindex’The page served a noindex directiveWebsite
Soft 404A not-found message returned with a 200 codeWebsite
Blocked due to unauthorized request (401)The page demanded credentialsWebsite
Not found (404)The URL returned a 404Website
Blocked due to access forbidden (403)The server returned a 403Website
Blocked due to other 4xx issueSome other 4xx responseWebsite

Two of these deserve extra attention because they are so often self-inflicted. A soft 404 means you are returning a friendly “nothing here” page with a 200 status — the engine sees success and thin content rather than absence, so tell the truth with a 404 or 410. And a 401 or 403 on pages you meant to be public usually means a staging rule or a firewall shipped to production.

The statuses that look like errors and are not

These are the ones that generate the most wasted work. Every one of them is Google reporting a decision it made, not a fault you introduced. Chasing them consumes the attention the real problems need.

StatusWhat it actually means
Alternate page with proper canonical tagConsolidation working exactly as intended
Duplicate without user-selected canonicalYou declared no canonical, so Google picked one
Duplicate, Google chose different canonicalYou declared one, Google disagreed
Page with redirectA non-canonical URL redirecting elsewhere, as designed
Discovered — currently not indexedKnown but not yet fetched
Crawled — currently not indexedFetched, assessed, and set aside

The two duplicate statuses are worth separating. “Duplicate without user-selected canonical” means you never declared a canonical and Google chose for you — usually harmless, but you have handed over a decision you could make yourself. “Google chose different canonical” means you did declare one and were overruled, which is a signal that the page you nominated is not the one the engine considers primary.

Reading “crawled, currently not indexed” honestly

This status means Google crawled the page but did not index it at that time. It does not identify one universal cause, and the live URL Inspection test cannot predict indexing success with certainty. There is no value in inventing a technical error that the report did not name.

On template-driven sites, useful things to review include near-duplication, canonical signals, internal links, and whether each page adds information that its siblings do not. Those are diagnostic checks, not a claim that every URL in this state failed a quality threshold.

  1. Open the URL Inspection tool and confirm the page returns 200, is self-canonical, and renders its main content.
  2. Compare the page against its three closest siblings. If you can swap the substituted values and get an equivalent page, the engine can see that too.
  3. Check whether anything internal links to it with descriptive anchor text, or whether it exists only in the sitemap.
  4. Ask what this page answers that no sibling answers. If there is no answer, that is the finding.
  5. Change the page so the answer exists, then request indexing — in that order.

“Discovered, currently not indexed” is a different signal

This one means the URL is known but has not been fetched. On a small site that is usually a short-lived state. At scale it is the clearest symptom Google names of a real crawl-budget constraint: too many URLs, too little demand, and the queue never reaches them.

If a large share of your URLs sit here, the productive work is reducing how many low-value URLs you ask the engine to consider — consolidating duplicates, returning proper 404s for what is gone, cutting redirect chains — rather than resubmitting individual pages.

Working the report instead of the symptoms

The report is organised by cause, which makes it tempting to work top-down by volume. Resist that. A thousand URLs under “Page with redirect” is usually a working redirect map, while forty under “Soft 404” is a genuine defect affecting real pages.

  • Fix everything in the error table first — those are unambiguous and mechanical.
  • Dismiss the informational statuses unless the volume contradicts your intent, such as canonical consolidation you never designed.
  • Treat “crawled, currently not indexed” as an editorial backlog, not a technical one.
  • Re-validate only after the underlying response or content has actually changed.

Try it on your own site

Questions and answers

Will resubmitting a URL force indexing?

No. Requesting indexing re-queues a crawl; it does not override the decision that followed the last one. Resubmit after the page has changed, not instead of changing it.

Why are some of my localized pages indexed and others not?

Templated pages that differ only by a substituted name read as near-duplicates. Engines commonly index a subset and drop the rest, and which ones survive tracks market demand more than markup. The fix is giving the dropped locales something of their own to say, not adjusting their tags.

Is “Alternate page with proper canonical tag” a problem?

No. It reports canonical consolidation working as designed. It only warrants a look if the volume is far larger than your canonical strategy would explain, which would suggest URLs are being generated that you did not intend.

What is the difference between soft 404 and 404?

A 404 tells the truth: nothing is here. A soft 404 shows a not-found message while returning a 200, so the engine records a successful fetch of a thin page. Return a real 404 or 410 for anything genuinely gone.

How long should I wait before treating non-indexing as a decision?

Compare against the same site rather than the clock. If comparable pages published at the same time were indexed within days and this one has sat for weeks, it has been assessed rather than delayed.