Skip to main content
← All guides

How search engines work

Three stages sit between a published page and a search result. Most SEO problems belong to exactly one of them, and naming the right one saves you from fixing the wrong thing.

The four stages, and why the distinction matters

A search engine does four separate things with your page, and it can succeed at one while failing the next. Discovery finds the URL. Crawling fetches it. Indexing decides whether to store it. Ranking decides, per query, whether to show it. Google's AI features are part of Search and still depend on the same crawl, index, and eligibility foundations, even though retrieval may fan out across related searches.

Almost every question that starts “why isn’t my page showing up” is really a question about which stage it stopped at. The answer changes what you do next completely, and the stages fail in ways that look similar from the outside — an absent page looks identical in the results whether it was never discovered, blocked at the fetch, judged not worth storing, or stored and simply outranked.

StageWhat happensHow it fails
DiscoveryThe URL becomes known, through a link or a sitemapNothing links to the page and no sitemap lists it
CrawlingThe URL is fetched and rendered like a browser wouldrobots.txt withholds permission, or the server errors out
IndexingThe engine decides whether to store the pageStored content is judged duplicate, thin, or not worth keeping
RankingStored pages are scored against a query's intentThe page is stored but a better answer exists for that query

Discovery: how a URL becomes known

Google states that the vast majority of new pages it finds each day arrive through links. A crawler already fetching one of your pages extracts the URLs in its markup and queues them. That makes internal linking the primary discovery mechanism on any site, not an optimisation applied afterwards.

Sitemaps are the secondary route. They matter most where linking is weakest: very large sites, brand-new sites with few inbound links, and sites heavy with media that has no natural anchor text. Google is explicit that listing a URL in a sitemap guarantees neither crawling nor indexing — it is a suggestion offered for consideration, not a queue that will be worked through.

The practical consequence is that a page nothing links to is relying on the weaker of the two mechanisms. A sitemap entry carries no context: it says a URL exists, but nothing about what the page is for or how it relates to anything else. A link carries both, through its placement and its anchor text.

Crawling: permission, capacity and demand

Crawling is bounded by two things Google names separately. The crawl capacity limit is how many simultaneous connections it will open without straining your server — it adapts to your response times and error rates. Crawl demand is how much it actually wants to fetch, driven by site size, update frequency, and its assessment of quality.

Google's own guidance is that most sites should not think about this at all. Crawl budget becomes a real concern on sites above roughly a million pages updating weekly, sites above ten thousand pages changing daily, or any site showing a large share of URLs stuck at “Discovered — currently not indexed”. Below that, a slow crawl is almost always a symptom of something else.

Where it does apply, the levers Google lists are unglamorous: consolidate duplicates, return a real 404 or 410 for pages that are gone rather than a soft 404, keep redirect chains short, improve server response time, and support conditional requests so unchanged pages can be answered cheaply. Note what is missing from that list — noindex is not a crawl-budget tool, because the crawler has to fetch the page to read the directive.

Rendering: the step people forget

Between fetching and indexing sits rendering. The engine executes the page roughly as a browser would, then indexes what the rendered output contains. Content that only appears after JavaScript runs can therefore be indexed — but it depends on that execution succeeding, and on the resources it needs being crawlable.

This is where a robots.txt rule can cause damage far from where it was written. Blocking a scripts or API path stops the crawler fetching the very files the page needs to build its content. The HTML is fetched fine, renders nearly empty, and gets indexed nearly empty. The status will not say “blocked” — the page itself was never blocked.

Indexing is a judgement, not a queue

The most common misreading of the whole pipeline is treating indexing as a waiting line: get crawled, wait your turn, get indexed. It does not work that way. After rendering, the engine decides whether the page is worth storing, and that decision can be no.

This is exactly what Search Console means by “Crawled — currently not indexed”. Nothing is broken, nothing is blocking, the fetch succeeded. The page was assessed and set aside. Near-duplicate pages generated from a template — differing only by a substituted name or number — are the usual reason, and resubmitting them changes nothing, because the input to the decision has not changed.

Ranking and Google's AI features share the same foundations

For a given query, Search retrieves and ranks relevant material. Google's AI Overviews and AI Mode can also use query fan-out, issuing multiple related searches across subtopics and data sources before generating a response with supporting links.

The consequence is simple: visibility in Google's AI Search features still depends on ordinary Search eligibility. Google says there are no additional technical requirements or special AI markup; crawling must be allowed, important content should be indexable and available in text, and preview controls such as nosnippet can limit what appears in AI features.

Try it on your own site

Questions and answers

Does blocking a page in robots.txt remove it from Google?

No. robots.txt stops the fetch, not the indexing. A blocked URL that other sites link to can still be indexed and listed without a description, because the crawler was never allowed to read the noindex that would have excluded it. To keep a page out of results, let it be crawled and serve a noindex directive, or require authentication.

How long should indexing take?

There is no guaranteed window. A new page can be indexed within hours or left for weeks, and Google states plainly that being crawled is no promise of inclusion. If a page has sat unindexed for weeks while comparable pages on the same site went in quickly, treat it as a judgement about the page rather than a delay.

Do I need to worry about crawl budget?

Almost certainly not. Google's thresholds are roughly a million pages updating weekly, or ten thousand pages changing daily. Below that, pages that are slow to appear are usually failing at indexing rather than waiting on crawl capacity.

Can JavaScript content be indexed?

Yes, because the engine renders pages before indexing them. The risk is not JavaScript itself but its dependencies: if robots.txt blocks the scripts or endpoints the page needs, it renders empty and is indexed empty, with no status anywhere naming the cause.