All posts
crawl budgettechnical seolog file analysis

Crawl Budget: When It Actually Matters (Usually It Does Not)

If your site has fewer than a few thousand URLs, crawl budget is not your problem, and anyone selling you crawl budget optimisation is selling you nothing.

Crawl Budget: When It Actually Matters (Usually It Does Not)

Crawl budget is the number of pages a search engine is willing to crawl on your site in a given period.

And for the overwhelming majority of websites, it does not matter at all.

If your site has a few thousand URLs or fewer, Google will crawl all of it, comfortably, and no amount of crawl budget optimisation will improve anything. Google has said as much repeatedly.

Crawl budget starts to matter at scale: hundreds of thousands of URLs, or a site that generates near-infinite URLs through faceted navigation, or a site so slow that the crawler gives up.

If a consultant is selling crawl budget optimisation to a 400-page website, they are selling you a solution to a problem you do not have.


When it genuinely does not matter

You have fewer than a few thousand URLs. Your site is not generating URL permutations from filters. Your server is not slow. And your Search Console shows Google crawling your pages regularly.

In that situation, crawl budget optimisation will change nothing, and the effort is better spent on literally anything else.

Your pages are not missing from the index because Google ran out of budget. They are missing because Google looked and decided against them, and those are completely different problems with completely different fixes. → Why Google will not index your pages

This distinction matters enormously, because "we have a crawl budget problem" is a much more comfortable thing to tell a client than "your content is not good enough."


When it genuinely does

A very large site. Hundreds of thousands of URLs, or millions. Ecommerce with a big catalogue. Publishers with decades of archive. Marketplaces. Classifieds.

Faceted navigation generating near-infinite URLs. This is the real one.

A shop with 10 filters, each with 5 options, can generate millions of URL combinations. ?colour=blue&size=9&brand=nike&sort=price&page=4

Google will happily crawl those millions of URLs, most of which nobody has ever searched for, while your actual category pages wait.

That is a genuine crawl budget problem and it is almost entirely a faceted navigation problem.Ecommerce SEO

A slow server. Google adjusts its crawl rate to what your server can handle. A slow site gets crawled less, and this is a real and under-appreciated cost of poor performance.Core Web Vitals

Enormous numbers of low-value URLs. Session IDs in URLs. Infinite calendars. Print versions. Every one of them is a page Google crawls instead of yours.


What to actually do, if you genuinely have the problem

Kill the URLs that should not exist.

This is 90% of the work. Not "optimise the crawl." Reduce the surface.

Faceted navigation: decide which filter combinations have genuine search demand, and make those real, indexable, linkable pages. noindex or block the rest. Nobody searches "blue size 9 Nike sorted by price ascending page 4."

Session IDs and tracking parameters: strip them, or canonicalise them away. → Canonical tags

Infinite spaces: calendars that generate a page for every date until 2087. They exist. We have found them.

Fix the redirect chains. Every hop is a crawl.

Fix the 404s that are being linked to internally. You are sending the crawler to dead ends, repeatedly, with your own links.

Make the server faster. A faster server gets crawled more.

Keep the sitemap clean and accurate. A sitemap full of redirects, 404s and noindex pages is actively misleading the crawler.XML sitemaps at scale

Link properly. Pages deep in the architecture with few internal links get crawled less, because your own site is telling Google they are unimportant.Internal linking at scale


Log file analysis, which is how you actually know

Everything above is inference. Log files are evidence.

Your server log records every single request, including every visit from every crawler.

Which means it tells you what Google is genuinely doing on your site, rather than what you assume it is doing.

What you find in there, and you will not find it anywhere else:

Which pages Googlebot actually crawls, and how often. Frequently a shock. Which pages it never visits at all. How much of your crawl is being spent on junk: parameters, redirects, 404s, filter permutations. On a big ecommerce site this number is regularly over half. Whether Googlebot is hitting errors you cannot see. How your crawl rate responds to server speed. And whether the AI crawlers are visiting at all, which is the newest and most useful thing in there. If OAI-SearchBot and PerplexityBot never appear in your logs, something is blocking them.The AI crawler guide

Log file analysis is unglamorous, it is genuinely difficult on a large site, and it is the only source of truth about crawling that exists.

And it is worth doing precisely once, at the start, on a large site, to find out whether you have a crawl problem at all before anybody sells you a solution to one.


What is crawl budget?

The number of pages a search engine is willing to crawl on your site in a given period. For the overwhelming majority of websites it does not matter at all.

Does crawl budget matter for my site?

Almost certainly not, if you have a few thousand URLs or fewer. Google will crawl all of it comfortably. Crawl budget matters at hundreds of thousands of URLs, with faceted navigation generating endless permutations, or on a very slow server.

My pages are not indexed. Is that a crawl budget problem?

Usually not. It is far more often that Google crawled the page, looked at it, and decided against indexing it. Those are completely different problems, and "we have a crawl budget issue" is a much more comfortable thing to tell a client than "the content is not good enough."

What causes real crawl budget problems?

Faceted navigation, overwhelmingly. Ten filters with five options each can generate millions of URL combinations, and Google will happily crawl them while your actual category pages wait.

How do I fix crawl budget?

Reduce the surface rather than optimising the crawl. Decide which filter combinations have genuine search demand, make those real pages, and block or `noindex` the rest. Then fix redirect chains, dead internal links, and server speed.

What is log file analysis?

Reading your server logs to see exactly what search engine crawlers actually do on your site, rather than what you assume. It is the only source of truth about crawling, and it will also tell you whether the AI crawlers are visiting at all.

Does site speed affect crawling?

Yes. Google adjusts its crawl rate to what your server can handle, so a slow site gets crawled less. It is a real and under-appreciated cost of poor performance.

Support Team

Still Have Questions on Your Mind?

Contact our support team and we’ll guide you every step of the way.

Get in touch

Want to work with a team that ships?

Twenty-minute call. No deck, no script. We'll tell you whether Stepwise fits your operation.

Book the fit call