Crawl Budget and Log File Analysis
Crawl budget is the number of URLs a search engine can and wants to crawl on a site in a given period. Log file analysis is the most accurate way to see how that budget is being spent.
  • Crawl budget is set by crawl capacity (what your server can handle) and crawl demand (how much Google wants your pages).
  • It is mainly a concern for very large or fast-changing sites.
  • The biggest wastes are faceted navigation, parameters, duplicates, soft 404s and redirect chains.
  • Server logs record every bot request and are the only complete source of crawl data.
  • Always verify bots by IP or reverse DNS, since user agents are easily faked.

Search engines do not have unlimited resources, so they ration how much they crawl on each site. Most small and medium sites never hit that limit. For sites with hundreds of thousands of URLs, heavy filtering or frequently changing inventory, crawl efficiency can decide whether new pages are found in hours or weeks.

What determines crawl budget

  • Crawl capacity limit: how many simultaneous connections Googlebot can use without slowing your server. Fast, healthy responses raise it. Errors and slow responses lower it.
  • Crawl demand: how much Google wants to crawl, driven by popularity, how often pages change and how many URLs it knows about.

Crawl budget is the overlap between the two. Crucially, low-quality or duplicate URLs soak up demand that could have gone to pages you care about.

Who should worry

Google's own guidance points to sites with over a million pages that change weekly, sites with over ten thousand pages that change daily, and sites with a large share of URLs stuck at "Discovered, currently not indexed". If that is not you, focus on quality and internal linking instead.

Where crawl budget goes missing

  • Faceted navigation and parameter combinations.
  • Session IDs and tracking parameters.
  • Duplicate content across protocols, hostnames and paths.
  • Soft 404s and thin pages.
  • Long redirect chains.
  • Infinite spaces such as calendars with endless next links.
  • Hacked or auto-generated pages.

How to improve crawl efficiency

  • Block truly valueless URL spaces in robots.txt.
  • Return 404 or 410 for removed pages so Google stops requesting them.
  • Fix redirect chains and internal links to redirected URLs.
  • Keep sitemaps clean and lastmod accurate.
  • Improve server response time.
  • Strengthen internal linking to important and new pages.
  • Support conditional requests so unchanged pages can return 304.

Noindex does not save crawl budget in the short term, because the page must be crawled for the tag to be read. Canonicals do not prevent crawling either.

Log file analysis

Every request to your server is recorded in an access log: timestamp, IP address, method, URL, status code, bytes sent, referrer and user agent.

 66.249.66.1 - - [07/Oct/2026:09:14:02 +0000] "GET /guides/canonical-tags HTTP/1.1" 200 18452 "-" "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" 

Crawling tools tell you what could be crawled. Logs tell you what was. If you use a CDN, get the logs from there, as many requests never reach your origin server.

Questions logs can answer

  • Which sections get the most and least crawl attention?
  • How often are key pages crawled?
  • What share of requests goes to parameter URLs, redirects or errors?
  • Are there pages in the sitemap that bots never request?
  • Are bots requesting orphan pages that are not in your crawl?
  • How quickly are new pages discovered?
  • Did crawl behaviour change after a release or migration?

Verify the bots

Anyone can claim to be Googlebot. Confirm with a reverse DNS lookup or by matching against the IP ranges that Google, Bing, OpenAI and others publish.

Search Console Crawl stats

The Crawl stats report (under Settings) is a good free alternative. It shows total requests, download size, average response time, host status, and breakdowns by response code, file type, purpose (discovery or refresh) and Googlebot type. It covers Google only and samples example URLs.

Common mistakes

  • Obsessing over crawl budget on a five-hundred-page site.
  • Using noindex to solve a crawl problem.
  • Blocking URLs in robots.txt that are already indexed and expecting them to vanish.
  • Analysing origin logs only, when the CDN serves most traffic.
  • Trusting user agent strings without verification.

The GEO angle

Logs are currently the best window into AI activity on your site. You can see which AI crawlers visit, how often, which pages they request and what status codes they get. User-triggered agents such as ChatGPT-User show real-time retrieval: each hit is a person asking a question for which your page was fetched. That is demand data you cannot get from any other source.

Add your title here

This is a paragraph. Writing in paragraphs lets visitors find what they are looking for quickly and easily. Make sure the title suits the content of this text.

Contact Us Amy Time