The ad platform tax: what owning your traffic actually takes
Every year the same buyer costs more to rent. How organic visibility compounds into an asset you own, where the compounding comes from, and…
SEO 7 min read
Large catalogs waste crawl budget on filter combinations and duplicate variants — leaving deep product pages undiscovered. Here’s how to find and fix it.
The short version
A store with ten thousand products assumes every one of them is discoverable. In practice, search engines spend a
limited amount of time crawling any given site, and on a large catalog that budget runs out long before it reaches
the pages furthest from the homepage.
The products that suffer aren’t random. They’re usually the newest additions, the lower-traffic variants, and
anything buried three or four clicks deep in a category structure — exactly the pages a long-tail search is most
likely to be looking for.
Faceted navigation is the biggest offender on most platforms. A collection with three filters and six options each
can generate hundreds of crawlable URL combinations from a single page — colour, size, price range, sort order —
and search engines will happily crawl every one of them unless told not to.
Search Console’s crawl stats report is the starting point — it shows what proportion of recent crawl activity hit
which parts of the site. Cross-referenced against a full site crawl, the pattern usually becomes obvious fast: a
small number of URL patterns eating a disproportionate share of attention.
Every unit of crawl spent on a filter combination is a unit not spent on a page that sells.
Adlux growth team
Crawl waste becomes noticeable once a catalog has a few thousand SKUs or heavy faceted filtering. Below that, other issues — like weak internal linking — usually matter more.
It usually shows up first as previously-unindexed pages getting crawled and indexed, which takes weeks rather than days. Ranking gains follow once those pages have had time to earn signal.
Noindex still requires a crawl to be seen, so it doesn’t save crawl budget the way a robots.txt disallow does. It’s the right tool for pages you want to keep accessible to users but out of the index — not for closing off crawl waste.
Crawl budget isn’t a ranking factor you optimise for its own sake — it’s a resource, and large catalogs spend most
of theirs on pages that were never going to convert. Redirecting it toward the products that actually sell is one
of the few SEO fixes that doesn’t require writing a single new page.
Written by the Adlux team from live account work. If you want this thinking applied to your catalog, the audit is free and the findings come in writing.