Crawl budget is the number of URLs Google is able and willing to crawl on your site. Google’s own definition is “the set of URLs that Google can and wants to crawl,” and it comes from two things: how much your server can handle, and how much Google wants your content. For most sites it doesn’t matter. Google’s guide is written for sites with about 1 million+ pages changing weekly, or 10,000+ pages changing daily.
So my honest answer for a typical small business or blog is no, you don’t need to worry about crawl budget. But there are a few situations where it becomes the real reason your pages aren’t getting indexed, and those are worth knowing.
This post walks through how to tell which group you’re in, what wastes crawling, and what to do about it.
Does Your Site Even Need to Worry About Crawl Budget?
Start here, because most crawl budget work I see is wasted effort on sites that never had a crawl problem. Google’s crawl budget management guide says it’s an advanced guide for three kinds of sites:
- Large sites (1 million+ unique pages) with content that changes about once a week
- Medium or larger sites (10,000+ unique pages) with content that changes daily
- Sites where a large share of URLs sit in Search Console as “Discovered – currently not indexed”
Google adds that these numbers are rough estimates, not exact thresholds. It also gives the simplest test I know: if your pages seem to be crawled the same day they’re published, you don’t need the guide.
Here’s the condition map I use.
| Your situation | Crawl budget a real issue? | What I’d do instead |
|---|---|---|
| Under a few thousand URLs, new posts indexed quickly | No | Work on content and links |
| Small site, pages “Discovered – currently not indexed” | Rarely | Check quality and internal links first |
| Store with filters creating thousands of parameter URLs | Often | Control faceted URLs, duplicates and soft 404s |
| 10,000+ pages changing daily (news, listings, inventory) | Yes | Full crawl budget work: logs, Crawl Stats, server |
| Any size, server slow or throwing 5xx errors | Yes, on the capacity side | Fix hosting and response times |
That small-site row matters. Gary Illyes wrote on the Google Search Central blog back in 2017 that a site with fewer than a few thousand URLs will be crawled efficiently most of the time. If a small site has indexing trouble, I look at page quality and internal links long before crawl budget. My guide to Discovered – currently not indexed covers that path.
How Does Google Decide How Much to Crawl?
Two forces set your crawl budget. Understanding which one is limiting you tells you which fix to reach for.
Crawl capacity limit is the server side. Google calls it hostload. It caps how much time your server spends holding connections open for Google, counting both parallel connections and how long they take. Google says every site starts with the same conservative default. The limit rises if your site responds consistently and response times stay stable or improve. It drops if the site slows down or returns 5xx errors or 429 rate-limit responses.
Crawl demand is the content side. Google lists three factors you can influence:
- Perceived inventory. Without guidance, Google tries to crawl most URLs it knows about. Google calls this the factor you can control the most.
- Popularity. More popular URLs get crawled more often.
- Staleness. Google recrawls to pick up changes. Site moves can trigger a temporary spike.
Two details in the current version of the guide surprise people. First, crawl budget is set per hostname, so www.example.com and shop.example.com each have their own. Second, the capacity limit is shared across all of Google’s crawlers. Heavy demand from AdsBot or Google Shopping can leave less room for Googlebot.
And if demand is low, Google crawls less even when your server could handle more. That’s why a faster server alone rarely fixes a site Google simply isn’t interested in.
How Do You Check Your Crawling in Search Console?
When I audit a site for crawling, I look at the data before touching anything. The Crawl Stats report lives under Settings in Search Console and covers the last 90 days. It only works for root-level properties, meaning a Domain property or a URL-prefix property at the root.
The report shows total crawl requests, total download size and average response time. Below that, it breaks requests down by response code, file type, crawl purpose (discovery of new URLs versus refresh of known ones) and Googlebot type. There’s also a host status panel covering robots.txt fetching, DNS resolution and server connectivity.
Here’s what I check, in order:
- Host status. If any of the three panels show recent problems, that’s a capacity issue. Fix it before anything else.
- Average response time. A rising line usually comes with falling crawl requests. That’s the capacity limit dropping in real time.
- By response. A large share of 301s, 404s or 5xx tells me Google spends time on URLs that don’t return content.
- By purpose. If almost everything is refresh and very little is discovery, new pages are struggling to get found.
- Example URLs. Click into each row. Parameter URLs, sorted versions and calendar pages show up here fast.
For big sites, I pair this with server log analysis. Logs show every Googlebot hit, not a sample, so you can see exactly which folders eat the crawling.
What Wastes Crawl Budget?
In my experience, most crawl waste comes from URLs nobody meant to create. These are the usual suspects on the sites I audit.
Faceted navigation and infinite spaces. Filters for size, color, price and sort order multiply into thousands of URL combinations, and I show how a store keeps them out of the crawl in my faceted navigation SEO guide. Google’s faceted navigation documentation calls this an “infinite URL space,” because crawlers can’t tell a useful filter URL from a useless one without crawling it first. Calendars with endless next-month links do the same thing.
Duplicate URLs. Tracking parameters, session IDs, printer versions and product variants that each get their own URL. Google’s guide says to consolidate duplicates so crawling goes to unique content, not unique URLs. I cover the store-specific version in duplicate content on ecommerce sites.
Soft 404s. Empty category pages and “no results” pages that return 200. Google’s guide says plainly that soft 404 pages “will continue to be crawled, and waste your budget.” My soft 404 guide shows how to find and fix them.
Redirect chains. Google’s guide says to avoid long redirect chains because they hurt crawling. Point internal links straight at the final URL. The 301 vs 302 redirects post covers chains in detail.
Slow pages. Google says that if it can load and render your pages faster, it may be able to read more of your site.
Should You Use Robots.txt or Noindex to Save Crawl?
This is where I see the most confusion, and the answer depends on what you want.
If you want Google to stop crawling a URL pattern, use robots.txt. Google’s guide says blocking with robots.txt prevents crawling, and gives infinite scrolling pages and differently sorted versions of the same page as examples. For filters, Google’s faceted navigation page recommends a robots.txt disallow on the filter parameters.
If you want a page out of the index, noindex is the right tool, but it doesn’t save crawl. Google’s guide says not to use noindex for crawl budget, because Google “will still request, but then drop the page” when it sees the tag, which wastes crawling time.
If a page is gone for good, return 404 or 410. Google says a 404 is a strong signal not to crawl that URL again, while URLs blocked in robots.txt stay in the crawl queue much longer.
Don’t use robots.txt to shuffle budget around. Google says blocking pages won’t move that crawling to other pages unless you’re already hitting your capacity limit. On a small blog, blocking every tag page won’t push Google toward your new posts.
Canonical tags sit in between. Google’s faceted navigation page says rel=”canonical” may decrease crawling of non-canonical versions over time, but that it’s less effective than robots.txt.
What If Googlebot Is Crawling Too Much?
Sometimes the problem runs the other way, and Googlebot hits a fragile server hard enough to slow it down for real visitors.
If it’s an emergency, Google’s reduce crawl rate guide says to return 500, 503 or 429 instead of 200 to crawl requests. Google treats these as a signal to slow down. Keep it short, though. The guide says a couple of hours or one to two days is fine, but if the same URLs return errors for multiple days, they may be dropped from Google’s index.
If it keeps happening, the real fix is capacity. Add server resources, cache pages, or put a CDN in front of the site. Google’s guide notes that if the URL Inspection tool shows “Hostload exceeded,” adding server resources is one of the two ways to raise crawl budget.
A tempting shortcut is blocking Googlebot entirely in robots.txt during a traffic spike. Please don’t. It does far more damage than a short run of 503s.
What If Google Isn’t Crawling Enough?
If Crawl Stats shows low and falling requests and new pages wait weeks, work through this list:
- Speed up server response. Response time is the lever you control most directly on the capacity side.
- Support 304 Not Modified. Google’s guide says a 304 tells Google to reuse its cached copy, which saves your bandwidth and server resources.
- Keep your sitemap current with accurate
lastmoddates. My XML sitemap best practices cover what belongs in it. - Cut the inventory. Fewer junk URLs means more attention for the ones that matter.
- Improve the content. For Search, Google factors in popularity, overall user value, content uniqueness and serving capacity. Three of those four are about the pages, not the server.
Is Crawl Budget a Ranking Factor?
No. In the same 2017 Search Central post, Gary Illyes wrote that a higher crawl rate won’t necessarily lead to better positions, and that crawling is necessary to be in the results but “it’s not a ranking signal.”
Crawling matters only because a page Google never crawls can’t be indexed or ranked. So I treat crawl budget as an access problem, not a ranking lever.
Crawl Budget Decision Matrix
| If | And | Then |
|---|---|---|
| New pages get crawled the same day | Any site size | Ignore crawl budget |
| Under a few thousand URLs | Pages not indexed | Fix quality and internal links first |
| Crawl Stats shows host status errors | Response time rising | Fix server capacity before anything else |
| Many parameter or filter URLs in Crawl Stats | Store or listings site | Block filter patterns in robots.txt, consolidate duplicates |
| Lots of soft 404s or redirects in “By response” | Any size | Return real 404/410s, link straight to final URLs |
| Googlebot overloading the server | Short emergency | Return 503 or 429 for hours, not days |
When to Bring In Help
If you run a large store, a listings site or a news site and Crawl Stats looks messy, crawl budget work is worth doing properly. It usually means log analysis, URL pattern decisions and server changes that need a developer.
That’s part of what we do under technical SEO at Skyranko. I’d start with a free audit to check whether crawling is really your bottleneck, because on most sites it isn’t.
Frequently Asked Questions
What Is Crawl Budget in Simple Terms?
It’s how many pages Google will crawl on your site in a given period. Google defines it as the set of URLs it can and wants to crawl. “Can” depends on your server’s health. “Wants” depends on how many URLs you have, how popular they are and how often they change.
How Do I Know if I Have a Crawl Budget Problem?
Check whether new pages get crawled soon after publishing. If they do, you don’t have one. If many URLs sit in “Discovered – currently not indexed” and the Crawl Stats report shows lots of parameter URLs, redirects or server errors, crawl budget may be part of the problem.
Does Noindex Save Crawl Budget?
No. Google still has to crawl a page to see the noindex tag, so it keeps requesting it. To stop crawling, use robots.txt. To remove a page for good, return a 404 or 410.
Do Small Websites Need Crawl Budget Optimization?
Usually not. Google’s guide targets sites with roughly 10,000+ pages changing daily or 1 million+ pages changing weekly. A small site gets more from better content, internal links and a clean sitemap.
Last updated: October 2026 by Mizanur Rahman



