Crawling, Rendering and Indexing
Search engines discover URLs by crawling, process the page by rendering it, and store what they find by indexing it. A page that fails at any of these three stages cannot rank, and cannot be cited by AI search either.
  • Crawling, rendering, indexing and ranking are separate stages, and each one can fail independently.
  • Discovery comes from links and sitemaps. A URL nobody links to is a URL that is easy to miss.
  • Blocking crawling does not prevent indexing. Use noindex for that.
  • Google renders JavaScript. Most AI crawlers do not.
  • Not every crawled page gets indexed. Quality and duplication both play a part.

Before a page can appear in search results it has to be discovered, fetched, rendered and stored. These stages are usually described as crawling, rendering and indexing. Ranking only happens afterwards, so most technical SEO is really about removing obstacles in this pipeline.

Stage 1: discovery and crawling

Crawlers such as Googlebot and Bingbot build a queue of URLs from links on pages they already know, XML sitemaps, redirects and submissions through tools like Search Console or IndexNow. Each URL is then fetched with an HTTP request.

Before fetching, the crawler checks robots.txt to see whether it is allowed in. It then looks at the response: the status code, the headers and the HTML. How often a URL is revisited depends on how important the crawler thinks it is and how often it changes.

Stage 2: rendering

Modern pages often rely on JavaScript to build their content. Google handles this by passing pages to a rendering service that runs an up-to-date version of Chromium, executes the scripts and looks at the final DOM. Rendering is resource-hungry, so it can happen after the initial HTML has been processed, and content that only appears after a click, a scroll or a login will not be seen.

The safest position is simple: anything that matters (main content, links, titles, canonicals, structured data) should be in the HTML the server sends.

Stage 3: indexing

Indexing is where the search engine decides what the page is about and whether to keep it. During this stage it:

  • Extracts text, links, images and structured data.
  • Groups duplicate and near-duplicate pages and picks one canonical URL to represent the cluster.
  • Respects any noindex instruction.
  • Assesses quality. Thin, duplicated or low-value pages may be crawled and then left out of the index.

Being crawled is not the same as being indexed, and being indexed is not the same as ranking.

Stage 4: serving and ranking

When someone searches, the engine retrieves candidate documents from the index and orders them using relevance, quality, authority, freshness, location, language and many other signals. Increasingly the same retrieval step also feeds AI-generated answers.

What controls each stage

  • Crawling: robots.txt, internal links, XML sitemaps, server speed and status codes.
  • Rendering: how much of your content depends on JavaScript, and whether scripts and CSS are crawlable.
  • Indexing: meta robots, X-Robots-Tag, canonical tags, content quality and duplication.
  • Ranking: content, links, user experience and everything else.

Common mistakes

  • Blocking a URL in robots.txt and expecting it to drop out of the index. A blocked URL can still be indexed from links alone, and Google cannot see a noindex tag on a page it is not allowed to fetch.
  • Leaving a staging noindex in place at launch.
  • Orphan pages that exist in the CMS but have no internal links.
  • Relying on client-side rendering for main content and links.
  • Assuming a sitemap submission guarantees indexing. It helps discovery, nothing more.

How to test

  • URL Inspection in Search Console shows whether a URL is indexed, when it was last crawled, which canonical was chosen and the rendered HTML Google saw.
  • The Page indexing report groups URLs by the reason they are not indexed.
  • Crawl stats in Search Console show request volumes, response codes and server response times.
  • Server logs show what crawlers actually requested, which is the only source of truth.

The GEO angle

AI assistants have their own crawlers, and many also draw on Google or Bing indexes. The big difference is rendering: crawlers from OpenAI, Anthropic and Perplexity generally fetch HTML without running JavaScript. A page that Google can read perfectly well may look empty to an AI crawler if its content is rendered client-side.

Add your title here

This is a paragraph. Writing in paragraphs lets visitors find what they are looking for quickly and easily. Make sure the title suits the content of this text.

Contact Us Amy Time