What Is Crawl Optimisation at Scale?
The objective is not more crawling. It is crawling of the right things. Most sites with a crawling problem are having their attention wasted rather than withheld. That reframing turns a technical curiosity into a decision a business owner can actually make.
It Is An Allocation Problem
This subject is usually discussed as though a site has a quantity of attention and the job is to obtain more of it. That framing is unhelpful, because the quantity is not the variable you control. Where it goes is.
What is actually happening. Distribution, not scarcity.
A site receiving substantial attention can still have its important pages visited rarely, because the attention is going elsewhere. The total can look entirely healthy while the distribution is wrong.
Why nobody chose that distribution. It emerges.
Attention follows links and links follow structure. Nobody decided that a filtered listing should receive more attention than a service page. It happened because the structure produces addresses faster than anybody notices.
Why the quantity framing leads nowhere. There is no dial.
You cannot request more, buy more or be granted more. Treating the problem as a shortage produces a conversation with no available action, which is why the subject reads as a curiosity when framed that way.
What the allocation framing gives you. Decisions.
Which addresses should exist, which should not, then which should be easy to reach. Those are answerable by somebody who understands the business rather than only by somebody who understands servers.
Where The Waste Comes From
Six recurring sources. Each is a category to recognise rather than a fix to apply, because the fixes are decisions and those come later.
Filtered and faceted listings. The largest by a distance.
Filtering controls that combine to produce enormous numbers of addresses. This has its own block because it dominates everything else on the list.
Parameters. Tracking, sorting, session identifiers.
Anything appended to an address creates a variant of the same page. Marketing campaigns and internal tools generate these routinely. Nobody records that they did.
Pagination. Long sequences of near-identical listings.
A category with many pages of results produces many addresses that differ only in which items appear on them.
Near-duplicate pages. Variants of one thing.
Product variations, printer versions, location pages differing only in a name, plus the same item reachable through several category routes.
Old redirects. Chains nobody removed.
Every redirect still being followed consumes attention on a page that does not exist any more.
Pages that no longer exist. Still being requested.
Addresses removed years ago continue to be visited for a long time afterwards, particularly where something still links to them.
Filters Are The Biggest Single Cause
A listing with several filters generates an enormous number of addresses. It happens without anybody deciding it should. This one cause routinely outweighs everything else combined.
Why the numbers grow so fast. Combinations multiply.
Each filter multiplies the possibilities rather than adding to them. A handful of filters with a few options each produces thousands of combinations. Adding one more filter multiplies that again. Nobody creating the sixth filter is thinking about the arithmetic.
Why it is invisible from inside. The pages were never made.
Nobody created those addresses. There is no list of them, they do not appear in a content management system and no one authored them. They exist because the combination is possible.
Who this affects most. Two sectors especially.
Retail and property, where filtering by several attributes is the entire point of the interface. Our ecommerce material covers the commercial side of that.
Why not simply remove the filters. They are the product.
Customers need them and removing them would damage the site as a shop. The question is not whether to have filters but which of the resulting addresses should be treated as pages worth attention, which is a decision rather than a technical measure.
What a sensible answer looks like. A small deliberate set.
A limited number of filter combinations that people genuinely search for, treated as real pages, with the rest working perfectly for customers while not being treated as destinations.
It Only Matters Above A Certain Size
A small site does not have this problem. Everything on it gets visited, allocation is not a meaningful concept and being sold crawl optimisation for a thirty page site is being sold nothing at all.
Why size is what makes it real. Allocation requires competition.
The problem only exists where there is not enough attention to go everywhere. Below that point the question does not arise, because everything is reached regardless of how the site is arranged.
Address count, not page count. The distinction that matters.
A site with a few hundred products and several filters may have vastly more addresses than a site with a few thousand static pages. Judging by what is in the content management system understates it badly.
The signal that it might apply to you. Generated addresses.
If your site has filtering, sorting, pagination or parameters, the number of addresses is much larger than the number of things you publish. That gap is what makes this relevant rather than the size of your catalogue.
Why we say this rather than sell it. The recommendation has to mean something.
An agency recommending this to every client is not making a recommendation. Saying plainly that most sites do not need it is what gives the advice weight when a site genuinely does.
How You Know You Have It
Three indications. Any one is worth investigating. All three together is close to conclusive.
Important pages rarely fetched. The clearest signal.
Your commercially significant pages receiving attention infrequently while other parts of the site are visited constantly. This is only visible in log file analysis, which is why the two subjects belong together.
Large numbers of addresses that should not exist. Discovered rather than known.
Finding that your site has vastly more addresses than pages, most of them combinations nobody authored. Businesses are consistently surprised by the scale of this.
A gap between published and indexed. The commonest first symptom.
Substantially more published than appears anywhere, without an obvious reason. That gap has several possible causes and this is one of the likelier ones on a large site. Our Search Console material covers reading that reporting.
What is not a signal. A general sense of slowness.
New pages taking a while to appear is normal and is not by itself evidence of an allocation problem. Diagnosing on that alone leads to work that changes nothing.
What to do before acting. Confirm rather than assume.
Every fix below is structural and therefore expensive to reverse. Establishing that the problem is real comes first.
The Fixes Are Decisions, Not Settings
People arrive at this subject expecting configuration. The answers are architecture. The reason most crawl work fails is that it was treated as the first when it needed to be the second.
Reduce what gets generated. The largest lever.
Deciding which filter combinations are permitted to become addresses at all. That is a product and interface decision rather than a technical one. It needs somebody who understands what customers actually filter by.
Decide what should not exist. A judgement about the catalogue.
Which pages are genuinely worth having, which are variants of something else and which were created for a reason that ended. Nobody can answer that from outside the business.
Make the important things easy to reach. Structure rather than instruction.
Reducing the number of steps between the front of the site and the pages that matter. That is information architecture rather than configuration. It is also the most durable of the three.
Why configuration alone disappoints. It suppresses symptoms.
Technical measures can hide addresses without reducing how many exist or improving what is reachable. The underlying arrangement is unchanged and the problem returns as the catalogue grows.
What that means for who does this. Not only a developer.
Every decision above requires somebody who knows the commercial value of the pages. A purely technical engagement produces technically correct changes to the wrong things.
Removing Pages Is Usually The Answer
Most large sites would perform better with substantially fewer pages. That is unwelcome and consistently true. It is the conclusion this subject arrives at from the allocation side.
Why fewer works. The arithmetic is direct.
Attention divided across fewer addresses means more of it reaching each. Nothing else on this page has that property, which is why removal outperforms every technical measure.
What makes it hard to agree to. Loss aversion.
Removing pages feels like removing potential. Every page might rank for something, so giving one up feels like closing a door rather than concentrating effort.
Why that instinct is wrong at scale. The pages compete with each other.
Weak pages do not sit harmlessly alongside strong ones. They consume attention that would otherwise reach the strong ones and they dilute what the site demonstrates about itself.
What removal actually means here. Rarely deletion.
Improve, consolidate or remove, in that order, which is covered properly in what is content pruning in SEO. Consolidation preserves value in a way deletion does not.
The agreement worth reaching first. Before anything is touched.
That the objective is fewer, better pages rather than more coverage. Without that agreed, the work stalls at the first page somebody defends.
Do Not Do This During Anything Else
Changing crawl behaviour alongside a migration or a redesign makes both impossible to read. This is the same discipline our site migrations material insists on, applied here.
Why it destroys the reading. Two causes, one effect.
Both kinds of change move the same numbers. Done together, no observation afterwards can be attributed to either, so you have made two significant changes and learned nothing about either.
Why it happens anyway. It looks efficient.
A rebuild is when everything is open, so bundling structural work into it seems sensible and saves a separate project. The saving is real and the cost is that nobody can tell what worked.
What it costs when it goes wrong. The diagnosis.
If performance falls afterwards, you cannot establish whether the migration or the crawl work caused it, so the response is guesswork and frequently involves undoing something that was helping.
The sequence that works. One, settle, then the other.
Complete the first change, allow performance to stabilise enough to read, then make the second. Slower on paper and considerably faster in practice, because nothing has to be unpicked.
If they genuinely cannot be separated. Record everything.
Where a deadline forces both at once, document precisely what changed and when, so that afterwards there is at least a record to reason from. The full series is on the advanced SEO guide.
Fewer, better
pages. Not
more coverage.
The decisions here need somebody who knows what your pages are commercially worth, not only somebody who knows servers. We establish whether the problem is real before proposing anything. We will tell you plainly if your site is not large enough to have it.
What the national tier covers:
The national tier is £1,550 a month. It suits businesses competing beyond one town. We will say so if a lower tier would serve you better.
Every guide.
One practice.
Semantic search, machine learning, topical authority, programmatic production, rendering, log files, crawl allocation, information architecture, content pruning and how strategy changes with scale.