bhavy@adluxmarketing.com +91 84605 70600
Back to blog

SEO 7 min read

Crawl budget on large catalogs: where deep product pages go to hide

Large catalogs waste crawl budget on filter combinations and duplicate variants — leaving deep product pages undiscovered. Here’s how to find and fix it.

The short version

  • Search engines allocate a finite amount of crawling to every site — crawl budget.
  • Large catalogs waste it on filter combinations, duplicate variants and thin pages.
  • The pages that lose out are usually the deep product pages that would actually convert.
  • Fixing crawl waste rarely requires new content — it requires stopping what's already there from hiding itself.

A store with ten thousand products assumes every one of them is discoverable. In practice, search engines spend a
limited amount of time crawling any given site, and on a large catalog that budget runs out long before it reaches
the pages furthest from the homepage.

The products that suffer aren’t random. They’re usually the newest additions, the lower-traffic variants, and
anything buried three or four clicks deep in a category structure — exactly the pages a long-tail search is most
likely to be looking for.

Where crawl budget actually goes #

Faceted navigation is the biggest offender on most platforms. A collection with three filters and six options each
can generate hundreds of crawlable URL combinations from a single page — colour, size, price range, sort order —
and search engines will happily crawl every one of them unless told not to.

  • Faceted filter URLs — near-infinite combinations from a handful of filters
  • Variant duplication — the same product live at a dozen colour and size URLs
  • Redirect chains — old category links hopping through two or three hops before landing
  • Thin auto-generated tag pages — one or two products, no real reason to exist independently

Finding out where your crawl is actually going #

Search Console’s crawl stats report is the starting point — it shows what proportion of recent crawl activity hit
which parts of the site. Cross-referenced against a full site crawl, the pattern usually becomes obvious fast: a
small number of URL patterns eating a disproportionate share of attention.

Every unit of crawl spent on a filter combination is a unit not spent on a page that sells.

Adlux growth team

The order that actually fixes it #

  1. Disallow the low-value filter parameters in robots.txt, while keeping pagination crawlable.
  2. Canonical every variant to its parent product page.
  3. Repoint internal links so nothing routes through a redirect chain.
  4. Surface orphaned products through collections and related-product blocks so they’re actually linked to.

Frequently asked questions #

Crawl waste becomes noticeable once a catalog has a few thousand SKUs or heavy faceted filtering. Below that, other issues — like weak internal linking — usually matter more.

It usually shows up first as previously-unindexed pages getting crawled and indexed, which takes weeks rather than days. Ranking gains follow once those pages have had time to earn signal.

Noindex still requires a crawl to be seen, so it doesn’t save crawl budget the way a robots.txt disallow does. It’s the right tool for pages you want to keep accessible to users but out of the index — not for closing off crawl waste.

The takeaway #

Crawl budget isn’t a ranking factor you optimise for its own sake — it’s a resource, and large catalogs spend most
of theirs on pages that were never going to convert. Redirecting it toward the products that actually sell is one
of the few SEO fixes that doesn’t require writing a single new page.

Written by the Adlux team from live account work. If you want this thinking applied to your catalog, the audit is free and the findings come in writing.

Written by

AdluxMarketing

Website

More articles by this author