Start with the status, not a diagnosis
The Page indexing report includes both genuine technical problems and expected exclusions. A robots.txt rule, noindex directive, authentication requirement or failing response can prevent indexing; duplicate and canonical statuses may be working exactly as intended; and a discovered URL may simply not have been crawled yet.
The most misleading state is “Crawled — currently not indexed”. It confirms that Google fetched the page but did not add it to the index at that time. Google explicitly says the page may or may not be indexed later and that there is no need to resubmit it just because of this status.
The statuses that are genuinely errors
These describe something on your side that stopped the page being stored. Google labels the source of each as Website when you can fix it yourself, and Google when the cause sits on their side. Correcting the underlying response is the whole fix.
| Status | What it means | Source |
|---|---|---|
| Server error (5xx) | The server returned a 500-level error | |
| Redirect error | A redirect chain looped or otherwise failed | Website |
| URL blocked by robots.txt | robots.txt withheld permission to fetch | Website |
| URL marked ‘noindex’ | The page served a noindex directive | Website |
| Soft 404 | A not-found message returned with a 200 code | Website |
| Blocked due to unauthorized request (401) | The page demanded credentials | Website |
| Not found (404) | The URL returned a 404 | Website |
| Blocked due to access forbidden (403) | The server returned a 403 | Website |
| Blocked due to other 4xx issue | Some other 4xx response | Website |
Two of these deserve extra attention because they are so often self-inflicted. A soft 404 means you are returning a friendly “nothing here” page with a 200 status — the engine sees success and thin content rather than absence, so tell the truth with a 404 or 410. And a 401 or 403 on pages you meant to be public usually means a staging rule or a firewall shipped to production.
The statuses that look like errors and are not
These are the ones that generate the most wasted work. Every one of them is Google reporting a decision it made, not a fault you introduced. Chasing them consumes the attention the real problems need.
| Status | What it actually means |
|---|---|
| Alternate page with proper canonical tag | Consolidation working exactly as intended |
| Duplicate without user-selected canonical | You declared no canonical, so Google picked one |
| Duplicate, Google chose different canonical | You declared one, Google disagreed |
| Page with redirect | A non-canonical URL redirecting elsewhere, as designed |
| Discovered — currently not indexed | Known but not yet fetched |
| Crawled — currently not indexed | Fetched, assessed, and set aside |
The two duplicate statuses are worth separating. “Duplicate without user-selected canonical” means you never declared a canonical and Google chose for you — usually harmless, but you have handed over a decision you could make yourself. “Google chose different canonical” means you did declare one and were overruled, which is a signal that the page you nominated is not the one the engine considers primary.
Reading “crawled, currently not indexed” honestly
This status means Google crawled the page but did not index it at that time. It does not identify one universal cause, and the live URL Inspection test cannot predict indexing success with certainty. There is no value in inventing a technical error that the report did not name.
On template-driven sites, useful things to review include near-duplication, canonical signals, internal links, and whether each page adds information that its siblings do not. Those are diagnostic checks, not a claim that every URL in this state failed a quality threshold.
- Open the URL Inspection tool and confirm the page returns 200, is self-canonical, and renders its main content.
- Compare the page against its three closest siblings. If you can swap the substituted values and get an equivalent page, the engine can see that too.
- Check whether anything internal links to it with descriptive anchor text, or whether it exists only in the sitemap.
- Ask what this page answers that no sibling answers. If there is no answer, that is the finding.
- Change the page so the answer exists, then request indexing — in that order.
“Discovered, currently not indexed” is a different signal
This one means the URL is known but has not been fetched. On a small site that is usually a short-lived state. At scale it is the clearest symptom Google names of a real crawl-budget constraint: too many URLs, too little demand, and the queue never reaches them.
If a large share of your URLs sit here, the productive work is reducing how many low-value URLs you ask the engine to consider — consolidating duplicates, returning proper 404s for what is gone, cutting redirect chains — rather than resubmitting individual pages.
Working the report instead of the symptoms
The report is organised by cause, which makes it tempting to work top-down by volume. Resist that. A thousand URLs under “Page with redirect” is usually a working redirect map, while forty under “Soft 404” is a genuine defect affecting real pages.
- Fix everything in the error table first — those are unambiguous and mechanical.
- Dismiss the informational statuses unless the volume contradicts your intent, such as canonical consolidation you never designed.
- Treat “crawled, currently not indexed” as an editorial backlog, not a technical one.
- Re-validate only after the underlying response or content has actually changed.