Technical SEO · Guide

What Is Crawling in SEO and How Does It Work?

Three things have to happen to a page, in order: it has to be found, then stored, then shown. Almost every technical problem in this section is a failure at one of those three points. Working out which one comes before any attempt to fix it.

Updated: August 2026
Written by: Andrew Odgers, Managing Director
Reading time: 15 minutes
The sequence this whole section is organised around

Three Things Have To Happen In Order

Nothing can happen to a page that is not found. Nothing can be shown that is not stored. Nothing is stored that was not judged worth storing. Those three stages happen in that order and each depends on the one before it.

Found. Something has to reach it.

A page nobody links to, nobody lists and nothing is told about may never be visited. Being on your website is not the same as being reachable, which is where a surprising number of problems begin.

Stored. Being kept for consideration.

Having been fetched, a page may or may not be retained. That is a judgement made about the page rather than an automatic consequence of visiting it, which block four covers.

Shown. Appearing for something.

A stored page competes to be presented when somebody searches. Being stored is a prerequisite rather than a promise. A great many stored pages never appear for anything at all.

Why the order matters diagnostically. It saves the wrong fix.

A page that is not appearing might not be stored, might not have been found, else might be stored and simply uncompetitive. Those are three entirely different problems and only one of them is technical.

Why the third stage is not technical. A different discipline.

Being shown is decided by whether your page answers the search better than the alternatives. That is content, competition and relevance rather than anything in this section, which is why a technically flawless site can be invisible.

How to use this page. Identify the stage first.

Everything below sits at one of the three stages. Establishing which one has failed narrows the work considerably. It also prevents a business paying for a technical fix to a content problem.

The distinction everything else depends on

Crawling Against Indexing

Being visited and being stored are separate events. A page can be visited and not stored. Confusing the two produces the wrong diagnosis constantly.

What crawling is. Fetching the page.

Software requests the address, receives whatever the server returns and reads it. That is the whole event, happening without any decision about whether the page is worth keeping.

What indexing is. Deciding to keep it.

A separate judgement about whether the page is worth storing for future consideration. That decision can go either way. It can also be revisited later on a page that was previously stored.

Why the confusion is so common. The words get used interchangeably.

People say a page is crawled when they mean it appears, then say it is indexed when they mean it was visited. Both errors send somebody looking in the wrong place.

What each failure looks like. Different symptoms.

A page never visited has no record of having been fetched. A page visited and not stored has a record and no presence. Those are distinguishable, which is what makes the distinction practical rather than pedantic.

The commonest wrong conclusion. Assuming a blocked page is hidden.

Preventing a fetch and preventing an appearance are different operations, which robots.txt covers. Mixing them up is the single most consequential error in this subject.

Described plainly

What A Crawler Actually Is

Software that fetches pages. It follows links to find new ones and it has no view of your site beyond what it can reach, which explains most of its limitations.

What it does. Requests and reads.

Asks a server for an address, takes what comes back and looks at it for links to other addresses. Then it does the same again. There is nothing more sophisticated at the core of it.

What it cannot do. Guess.

It has no map of your site, no list of what should exist and no way of knowing about a page nothing points at. Your understanding of your own site is not available to it.

Why that matters practically. Reachability is everything.

A page linked from your navigation will be found. A page reachable only by typing the address, else only through a form, may not be. That is a structural property of your site rather than a setting.

How it behaves on a site. Politely and continuously.

It does not fetch everything at once, because that would overwhelm a server. It returns repeatedly over time, which is why changes are noticed later rather than immediately.

What it does with what it finds. Passes it on.

Fetching and judging are done by different parts of the system. The software collecting pages does not decide what happens to them, which is why a page can be fetched reliably and never stored.

What identifies it. A stated identity.

Well behaved crawlers announce what they are when requesting a page, which is how a server can treat them differently from a person. Not everything fetching your pages is well behaved.

With the qualification people miss

Being Stored

Being stored means the page has been kept so it can be considered when somebody searches. It does not mean the page will ever appear. A page can be stored and never shown for anything.

What storage involves. Keeping and understanding.

The content is retained along with an assessment of what the page is about. That is what allows it to be considered later without being fetched again each time.

Why storage is not appearance. Competition.

Every stored page competes with every other stored page for any given search. A page can be perfectly stored, accurately understood and simply less useful than the alternatives.

Why that matters commercially. It changes who fixes it.

A page not stored is frequently a technical problem. A page stored and never shown is almost always a content or competition problem. No amount of technical work addresses it.

What gets refused. Several things.

Pages judged near-duplicates of something already stored, pages with very little on them and pages whose content could not be retrieved properly. None of those is an error in the usual sense.

Storage is not permanent. The point people forget.

A page stored once can be dropped later, particularly where it is thin or duplicated. Losing presence without anything breaking is a normal outcome rather than a fault.

By what it answers rather than which buttons to press

Checking Whether Your Site Can Be Crawled

Four questions cover it. Knowing the questions is more durable than knowing an interface, because interfaces move and the questions have been stable for years.

Is anything being fetched at all. The floor.

Whether your site is being visited, plus roughly how often. A site receiving no visits from crawlers has a fundamental problem rather than an optimisation opportunity.

Is anything being refused. The instruction check.

Whether your own site is telling crawlers to stay away, deliberately or otherwise. This is where the most damaging single mistake in the cluster hides, which the robots pillar covers.

Which pages are stored. The comparison that matters.

How many of your pages are retained against how many you have. A large gap either way is informative: too few suggests a problem, while far more than you expected suggests addresses you did not know about.

What is failing when fetched. The error check.

Whether requests are succeeding or returning problems. Block six covers how to read those without treating every one as urgent.

Where to look. The reporting.

All four are answerable from the search console reporting for your site. How those reports work belongs in our Google Search Console material rather than here.

Most of them are not urgent

Crawl Errors

An error means a request did not succeed. Most reported errors do not need action. Knowing which do is worth more than the list itself.

The categories. Three broadly.

The page was not there, the server failed, else the site itself refused the request. Those have entirely different causes and entirely different levels of seriousness.

Which ones matter. Server failures.

Repeated failures where the server could not respond are the ones to act on, because they affect everything rather than one page. Status codes covers reading them.

Which ones usually do not. Missing pages.

A page that no longer exists reporting as missing is the system working correctly. A site with none of these has either never changed or is hiding something.

Why the reports look alarming. Counts without context.

Hundreds of items presented as errors reads as a crisis. Most of them are records of pages that were deliberately removed, which is why the count is a poor guide to urgency.

How to tell the difference. Two questions.

Does anything link to it, plus is anybody arriving at it. An error nobody encounters is a report entry rather than a problem. Prioritising by those two questions is most of the work.

The symptom rather than the fix

Crawl Traps

Something on the site generating endless addresses from a finite amount of content. The result is attention spent on addresses that lead nowhere useful. The symptom is recognisable before the cause is.

Where they come from. Four usual sources.

Filters that can be combined in any order, tracking added to addresses, calendars that will generate any date requested, then pagination with no end. Each turns a handful of pages into an unbounded number of addresses.

Why filters are the commonest. Combinations multiply.

A shop with several filters applied in any combination produces an enormous number of distinct addresses showing overlapping sets of the same products. Nobody built those pages and they all exist.

What the symptom looks like. More addresses than pages.

Reporting showing vastly more addresses than your site has actual pages. That gap is the signal. It is visible without knowing what is causing it.

Why calendars are the strangest case. They are genuinely infinite.

A calendar willing to display any month will produce a page for any date requested, indefinitely into the future. Something following links will keep going.

What to do about it. Establish the source first.

The fix depends entirely on what is generating them, which is a developer question. Identifying that a trap exists is the part a business owner can do.

The scoping that matters commercially

Crawl Budget

The idea that a site receives a finite amount of crawling attention. A real consideration on very large sites and largely irrelevant on small ones, which is the part nobody selling it mentions.

What the concept describes. Finite attention.

No site is fetched continuously without limit, so on a large enough site the order and frequency of fetching becomes a genuine constraint. That is the whole idea.

When it genuinely applies. At considerable scale.

Sites with very large numbers of pages, else with addresses being generated faster than they can be fetched. Large retailers, publishers and classified sites, which is a small proportion of businesses.

When it does not. Almost everywhere else.

A business with a modest number of pages is not constrained by crawling attention. There is nothing to optimise, because everything gets fetched comfortably within whatever attention the site receives.

Why we state that plainly. It gets sold.

A business being sold crawl budget work for a fifty page site is being sold nothing at all. The work will be done, a report will follow and no outcome is possible because there was no constraint.

The exception worth knowing. Traps change the arithmetic.

A small site with a trap generating endless addresses can have a crawling problem despite being small, which is a trap problem rather than a budget one and is fixed at the source.

Where the large site case is covered. Elsewhere.

Allocation at scale belongs in our advanced SEO material, which owns that argument properly. Nothing here states a figure, because no meaningful one exists.

The decision rather than the syntax

What Should Not Be Crawled

Some pages genuinely should not be fetched. The list is shorter than most people assume. Deciding which is a judgement. It is also a different question from what should not appear in results.

Internal functions. The clearest case.

Administrative areas, account pages, checkout steps and anything requiring a login. None of it is useful to anybody arriving from a search and none of it should be fetched.

Generated variations. The volume case.

Filtered and sorted versions of a listing, where the underlying page already exists in a canonical form. This is where trap prevention and this decision overlap.

Duplicated utility pages. The small case.

Print versions, internal search results and similar variants that exist for a purpose other than being found. Modest in number and worth handling deliberately.

What should not be on the list. Anything you want found.

Blocking is applied far too broadly in practice, usually by somebody being cautious. A page blocked by mistake is invisible. Nothing on the page indicates that anything is wrong.

The distinction that catches everybody. Blocking is not removing.

Preventing a fetch does not prevent an appearance. A blocked page can still be listed. The robots pillar covers why, plus why the wrong instrument is used constantly.

The strategic point

Not Everything Should Be Indexed

A business wanting every page stored has usually not asked which pages deserve to be. More is not the objective. Treating it as one produces a worse site rather than a bigger presence.

Why more is not better. Weak pages dilute.

A site where most stored pages are thin is judged partly on those pages. Removing or improving them frequently helps more than adding anything, which is counterintuitive to most owners.

What the usual candidates are. Four kinds.

Tag and category pages nobody reads, near-duplicate location pages, thin product variants and old posts nobody would defend. Every site accumulates these without deciding to.

The question worth asking. Would you send somebody here.

If you would not send a prospective customer to a page, it is not clear why you want it found. That single test resolves most of the list quickly.

What this connects to. Pruning.

Deciding what a site should contain rather than what it happens to contain is covered properly in our advanced SEO material, which owns the content pruning argument.

What it does not mean. Deleting things.

Keeping a page out of an index is not the same as removing it from the site. Plenty of pages should exist for customers and have no business being found by strangers, which is a normal state rather than a compromise.

What it costs to get wrong. Both directions.

Too much stored dilutes the site with pages nobody would defend. Too little means something you sell is invisible. Neither error announces itself, which is why the decision is worth taking deliberately rather than inheriting whatever a platform did.

Why it belongs on a crawling page. It is the same decision.

Every question in this section eventually becomes which pages you want found, stored and shown. Answering that deliberately is the whole of the strategic work here.

Website migrations

Which of the three
stages actually
failed?

Not found, not stored, else stored and uncompetitive. Those are three different problems and only one of them is technical. Establishing which one prevents a business paying for a technical fix to a content problem, which happens more than anybody admits.

What we look at:

What is being fetched What your site refuses Stored against published Errors worth acting on Address generation Blocking decisions Index scope A named owner

If your site is small, we will tell you crawl budget work would achieve nothing.

The full guide series

Every guide.
One practice.

How search finds and stores a site, what your own files are telling it, addresses and duplication, status codes, mobile, hosting and what a real audit contains.

Questions people ask

Crawling and Indexing, Briefly

What is the difference between crawling and indexing?
Crawling is fetching the page. Indexing is a separate judgement about whether to keep it. A page can be visited and not stored, which is why the distinction is practical rather than pedantic: a page never visited has no record of being fetched, while a page visited and not stored has a record and no presence. Those are distinguishable.
Our page is indexed but does not appear anywhere. Why?
Being stored is a prerequisite rather than a promise. Every stored page competes with every other for any given search, so a page can be perfectly stored, accurately understood and simply less useful than the alternatives. That is almost always a content or competition problem, which no amount of technical work addresses.
Search Console shows hundreds of crawl errors. How bad is that?
Usually not bad. Most are records of pages deliberately removed, which is the system working correctly, so the count is a poor guide to urgency. Ask two questions of each: does anything link to it, plus is anybody arriving at it. An error nobody encounters is a report entry rather than a problem. Repeated server failures are the ones to act on.
We have far more addresses than pages. What causes that?
Something generating endless addresses from finite content. Usually filters that combine in any order, tracking added to addresses, calendars that will produce any date requested, else pagination with no end. A shop with several filters applied in any combination creates an enormous number of addresses showing overlapping sets of the same products. Nobody built any of them.
Do we need crawl budget work?
Almost certainly not. It is a genuine constraint on very large sites and largely irrelevant on small ones, because everything gets fetched comfortably. A business being sold crawl budget work for a fifty page site is being sold nothing: the work will happen, a report will follow and no outcome is possible because there was no constraint. The exception is a site with a trap.
Should we try to get every page indexed?
No. Wanting to usually means nobody has asked which pages deserve it. A site where most stored pages are thin is judged partly on those pages, so removing or improving them frequently helps more than adding anything. The test is simple: if you would not send a prospective customer to a page, it is not clear why you want it found.