Site Migrations · Guide

How to Audit Your Current Site Before Migrating

This is the one stage of a migration that cannot be done retrospectively, which makes it both the most important and the most frequently skipped. Without it there is no way to know what was lost, no way to prove anything recovered and no way to separate a migration problem from an unrelated one.

Updated: August 2026
Written by: Andrew Odgers, Managing Director
Reading time: 10 minutes
Open with the one that matters

You Cannot Do This Afterwards

Once the old site is gone, the record of it is gone. Not difficult to obtain, not expensive to reconstruct. Gone. That single fact is what makes the audit the insurance policy for the entire project.

What disappears with it. Three things at once.

What every page said. Which addresses existed. How the site performed while it was still itself. None of those can be recovered from the replacement, because the replacement is the thing you would be measuring.

Why that matters in practice. Every later question needs it.

Did we lose anything. Has it come back. Was this the migration or something else. All three are unanswerable without a record taken beforehand. All three get asked.

Why it gets skipped. It produces nothing visible.

The audit delivers no page, no design and no improvement. It delivers a file that sits unopened unless something goes wrong, which makes it easy to defer and easy to cut.

The asymmetry worth stating. Small cost, unlimited downside.

Doing it costs a modest amount of time before a project that is already happening. Not doing it costs the ability to diagnose anything afterwards, permanently.

Five records, described as records

What The Audit Has To Capture

The audit is a set of records rather than a technique. What matters is that each exists somewhere findable, not how it was produced.

A complete list of addresses. Everything that resolves.

Not the pages in the menu and not what somebody would list from memory. Everything, including parameters, filters, pagination, tags and old addresses still reachable through existing redirects.

What each one ranks for. Attached to the address.

Recorded page by page rather than as a site total, because a total tells you nothing about which pages to check first afterwards.

Which have external links. The irreplaceable part, per block six.

Recorded specifically, because these are the addresses that must never be allowed to stop working.

Which have traffic. Not the same list.

Pages people arrive at, which overlaps with the ranking list without matching it.

What content sits on them. The words themselves.

Enough of each page to demonstrate later what it contained, which is the only defence against content quietly disappearing in a rebuild.

What recovery gets measured against

Benchmarks

A benchmark is a record of how the site performed before anything changed. Without one, the conversation after a migration is an exchange of opinions about whether things feel worse.

What to record. Four things.

Overall visibility. Traffic broken down by page rather than in total. Conversions or enquiries, being whatever the business actually counts. And the specific queries that matter commercially.

Why page level matters more than totals. Totals hide movement.

A site total can look stable while an important section has collapsed and something less valuable has risen to replace the numbers. Only a page level record reveals that.

What period to capture. Enough to see the pattern.

Long enough that seasonality is visible, because comparing a quiet month afterwards against a busy month before produces a decline that has nothing to do with the migration.

Where the figures come from. Block eight.

Described by what each source contributes rather than by name.

What a benchmark is not. A target.

It records what was, so that what follows can be compared to it. It does not establish what should happen next. Nobody should present it as a level performance will return to.

The commercially useful block

Not Every Page Deserves To Move

A migration is one of the few natural opportunities to remove pages that do nothing. Most sites accumulate them and almost nobody ever deletes anything deliberately.

What tends to be there. Years of residue.

Announcements about events that happened, pages built for campaigns that ended, near-duplicates created by somebody who could not find the original and drafts that were published by accident.

Why removing them helps. Two reasons.

Less to map, test and check, which shortens the project. And a site that says only what it means to say is easier to understand, for people and for anything assessing it.

The rule that protects you. Never on instinct.

A page is removable if it has no traffic, no external links and no purpose anybody can articulate. If it fails any of those three, it moves. Pages with links or traffic are never candidates, regardless of how they look.

How it should happen. Recorded in the map.

A deliberate decision written down, so that afterwards nobody wonders whether an address stopped working by accident. That distinction matters enormously during any later diagnosis.

Because you cannot check everything later

Find The Pages That Matter Most

After launch, a large site cannot be checked page by page. Knowing in advance which addresses carry the value determines what gets looked at first, which is the difference between finding a problem and finding it in week six.

What makes a page matter. Three overlapping things.

It carries external links. It brings traffic. It leads to something commercial. Most valuable pages score on more than one. A page scoring on all three is the first thing to verify after launch.

Why the answer surprises people. It is rarely the home page.

The pages carrying accumulated value are frequently unglamorous: an old guide, a specific product page, something written years ago that quietly attracted links. Nobody would nominate them from memory.

What to produce. A short priority list.

A modest number of addresses, ranked, that get checked first on launch day and watched hardest afterwards.

How it changes the mapping. These are mapped by hand.

Automated matching handles the bulk of a site. This list is the part that should never be automated, per URL mapping for a site migration.

The one thing that cannot be rebuilt

External Links Are The Irreplaceable Part

Content can be rewritten. Rankings can recover. Links pointing at addresses that stop working are lost in a way nothing else on a website is, which is why they are recorded separately.

Why they cannot be recreated. Somebody else owns them.

Every one of them was a decision made by another organisation, frequently years ago, frequently by somebody who has left. You cannot ask for them again and you certainly cannot ask at the scale a mature site has accumulated.

How they are lost. Quietly, by omission.

The linked address is missed from the map, so it stops working. Nothing announces this. The link still exists on the other site and now points at nothing, which is worse than not existing.

Where they usually sit. On old addresses.

Links accumulate over time, so the oldest addresses often carry the most. Those are also the addresses most likely to be forgotten, which is a bad combination and the reason block two insists on completeness.

What to do with the list. Treat it as the priority.

Every address on it maps by hand, gets tested individually before launch and gets checked again afterwards. Nothing else in the audit earns that treatment.

Protects everybody, including the developer

Record What Is Already Broken

Almost every site has existing problems. Recording them before the migration is what stops all of them being attributed to it afterwards.

What tends to be there already. Ordinary accumulation.

Pages that were already slow, addresses that were already returning errors, duplicated content nobody noticed and sections that stopped performing months earlier for unrelated reasons.

What happens without the record. Everything becomes the migration's fault.

After a launch, any problem anybody finds is assumed to have been introduced by it. That is a reasonable assumption and frequently wrong. Without a record there is no way to demonstrate otherwise.

Why this protects the build team too. It cuts both ways.

A developer who inherits a site with existing faults, holding no record of them, will be blamed for all of them. Establishing the starting condition is as much in their interest as anybody's.

What to do with what you find. Fix it or note it.

Some faults are worth correcting before the move so they are not carried across. Others are worth recording and leaving, so that the migration is not held responsible for them.

Sources by what they contribute

Where The Data Comes From

Four kinds of source contribute to an audit. None is complete on its own. That last point is the one worth taking away.

A crawl of the site. What exists and what is on it.

The most complete picture of addresses and content, produced by walking the site as anything else would. It finds what is reachable and it will not find addresses nothing links to.

The search engine's own reporting. What it has actually seen.

Which addresses are known, which queries produce impressions and which pages receive clicks. This finds addresses a crawl misses, including ones nobody links to internally.

Analytics. What people actually did.

Which pages received visits and what happened next. This is the only source that connects addresses to commercial outcomes.

A link source. Who points at what.

Which addresses have external links, per block six. This is the source most likely to be incomplete, because no provider sees every link.

Why more than one is needed. Each has blind spots.

An address with no internal links and no traffic can still carry external links. The union of the sources is the inventory. Any single one of them is a partial view presented confidently. The mechanism for the search and analytics sources is covered in our guides for each of those platforms.

Easier now than later

Keep The Old Site Somewhere

A copy or archive of the old site is worth having. It is trivial to arrange before launch and frequently impossible afterwards.

What it gives you. A reference.

The ability to look at what a page actually said, rather than at a record describing it. When somebody asks whether a section was shortened, this settles it in seconds.

What it is not. A rollback.

A copy is for reference rather than for restoring. Returning a site to a previous state after addresses have moved is a second migration, covered on the planning page.

Why it becomes impossible. The environment goes.

Hosting gets cancelled, the platform subscription lapses and access is lost with the handover. None of that is unreasonable and all of it happens quickly once the new site is live.

What to keep. The content, at minimum.

Even a stored copy of the pages is enormously more useful than nothing, at almost no cost to produce. The planning sequence is in how to plan a site migration, what to watch afterwards is in monitoring after migration. The full series is on the website migration guide.

Website migrations

Once it is gone,
the record
is gone.

Every address recorded rather than the pages somebody remembers, the benchmark taken page by page because totals conceal a collapsed section, the linked addresses listed separately because nothing else is irreplaceable, existing faults noted so the migration is not blamed for them, with a copy of the old site kept while that is still possible.

What a migration engagement covers:

Pre-migration audit Performance benchmark Address inventory URL mapping Redirect testing Launch day checks Notification Post-launch monitoring Recovery diagnosis

A migration is a one-off project rather than a monthly service, so it is scoped and quoted against the size of the site and how much is changing at once.

The full guide series

Every guide.
One practice.

What a migration is, the types, when to do it, planning, the checklist, auditing, mapping, redirects, internal links, tags, downtime, crawling, structured data, performance, notifying, monitoring, recovery and the mistakes.

Questions people ask

Auditing Before A Migration

Can we do the audit after launch if we run out of time?
No. This is the only stage where that is true. Once the old site is replaced, the record of what every page said, which addresses existed and how it performed is gone rather than merely difficult. It cannot be recovered from the replacement, because the replacement is the thing you would be measuring against.
Why does it get skipped so often?
Because it produces nothing visible. There is no page, no design and no improvement at the end of it, just a file that sits unopened unless something goes wrong. That makes it easy to defer and easy to cut. The asymmetry is worth stating plainly: modest cost before a project that is happening anyway, against permanently losing the ability to diagnose anything.
What should the benchmark contain?
Overall visibility, traffic broken down by page rather than in total, whatever the business actually counts as a result and the queries that matter commercially. Page level matters most, because a site total can look stable while an important section has collapsed and something less valuable has risen to replace the numbers.
Should we delete pages while we are migrating?
It is one of the few natural opportunities to, though never on instinct. A page is removable only if it has no traffic, no external links and no purpose anybody can articulate. Fail any of those three and it moves. Record every removal in the map as a deliberate decision, so that afterwards nobody wonders whether an address stopped working by accident.
Why record problems that already exist?
Because otherwise they all become the migration's fault. After a launch, any problem anybody finds is assumed to have been introduced by it, which is a reasonable assumption and frequently wrong. This protects the build team as much as anybody: a developer inheriting existing faults with no record of them will be blamed for all of them.
Is one data source enough?
No, because each has blind spots. A crawl finds what is reachable and misses addresses nothing links to. Search reporting finds addresses a crawl misses. Analytics is the only source connecting addresses to commercial outcomes. Link data is the most likely to be incomplete, since no provider sees every link. The union of them is the inventory.