How to Audit Your Current Site Before Migrating
This is the one stage of a migration that cannot be done retrospectively, which makes it both the most important and the most frequently skipped. Without it there is no way to know what was lost, no way to prove anything recovered and no way to separate a migration problem from an unrelated one.
You Cannot Do This Afterwards
Once the old site is gone, the record of it is gone. Not difficult to obtain, not expensive to reconstruct. Gone. That single fact is what makes the audit the insurance policy for the entire project.
What disappears with it. Three things at once.
What every page said. Which addresses existed. How the site performed while it was still itself. None of those can be recovered from the replacement, because the replacement is the thing you would be measuring.
Why that matters in practice. Every later question needs it.
Did we lose anything. Has it come back. Was this the migration or something else. All three are unanswerable without a record taken beforehand. All three get asked.
Why it gets skipped. It produces nothing visible.
The audit delivers no page, no design and no improvement. It delivers a file that sits unopened unless something goes wrong, which makes it easy to defer and easy to cut.
The asymmetry worth stating. Small cost, unlimited downside.
Doing it costs a modest amount of time before a project that is already happening. Not doing it costs the ability to diagnose anything afterwards, permanently.
What The Audit Has To Capture
The audit is a set of records rather than a technique. What matters is that each exists somewhere findable, not how it was produced.
A complete list of addresses. Everything that resolves.
Not the pages in the menu and not what somebody would list from memory. Everything, including parameters, filters, pagination, tags and old addresses still reachable through existing redirects.
What each one ranks for. Attached to the address.
Recorded page by page rather than as a site total, because a total tells you nothing about which pages to check first afterwards.
Which have external links. The irreplaceable part, per block six.
Recorded specifically, because these are the addresses that must never be allowed to stop working.
Which have traffic. Not the same list.
Pages people arrive at, which overlaps with the ranking list without matching it.
What content sits on them. The words themselves.
Enough of each page to demonstrate later what it contained, which is the only defence against content quietly disappearing in a rebuild.
Benchmarks
A benchmark is a record of how the site performed before anything changed. Without one, the conversation after a migration is an exchange of opinions about whether things feel worse.
What to record. Four things.
Overall visibility. Traffic broken down by page rather than in total. Conversions or enquiries, being whatever the business actually counts. And the specific queries that matter commercially.
Why page level matters more than totals. Totals hide movement.
A site total can look stable while an important section has collapsed and something less valuable has risen to replace the numbers. Only a page level record reveals that.
What period to capture. Enough to see the pattern.
Long enough that seasonality is visible, because comparing a quiet month afterwards against a busy month before produces a decline that has nothing to do with the migration.
Where the figures come from. Block eight.
Described by what each source contributes rather than by name.
What a benchmark is not. A target.
It records what was, so that what follows can be compared to it. It does not establish what should happen next. Nobody should present it as a level performance will return to.
Not Every Page Deserves To Move
A migration is one of the few natural opportunities to remove pages that do nothing. Most sites accumulate them and almost nobody ever deletes anything deliberately.
What tends to be there. Years of residue.
Announcements about events that happened, pages built for campaigns that ended, near-duplicates created by somebody who could not find the original and drafts that were published by accident.
Why removing them helps. Two reasons.
Less to map, test and check, which shortens the project. And a site that says only what it means to say is easier to understand, for people and for anything assessing it.
The rule that protects you. Never on instinct.
A page is removable if it has no traffic, no external links and no purpose anybody can articulate. If it fails any of those three, it moves. Pages with links or traffic are never candidates, regardless of how they look.
How it should happen. Recorded in the map.
A deliberate decision written down, so that afterwards nobody wonders whether an address stopped working by accident. That distinction matters enormously during any later diagnosis.
Find The Pages That Matter Most
After launch, a large site cannot be checked page by page. Knowing in advance which addresses carry the value determines what gets looked at first, which is the difference between finding a problem and finding it in week six.
What makes a page matter. Three overlapping things.
It carries external links. It brings traffic. It leads to something commercial. Most valuable pages score on more than one. A page scoring on all three is the first thing to verify after launch.
Why the answer surprises people. It is rarely the home page.
The pages carrying accumulated value are frequently unglamorous: an old guide, a specific product page, something written years ago that quietly attracted links. Nobody would nominate them from memory.
What to produce. A short priority list.
A modest number of addresses, ranked, that get checked first on launch day and watched hardest afterwards.
How it changes the mapping. These are mapped by hand.
Automated matching handles the bulk of a site. This list is the part that should never be automated, per URL mapping for a site migration.
External Links Are The Irreplaceable Part
Content can be rewritten. Rankings can recover. Links pointing at addresses that stop working are lost in a way nothing else on a website is, which is why they are recorded separately.
Why they cannot be recreated. Somebody else owns them.
Every one of them was a decision made by another organisation, frequently years ago, frequently by somebody who has left. You cannot ask for them again and you certainly cannot ask at the scale a mature site has accumulated.
How they are lost. Quietly, by omission.
The linked address is missed from the map, so it stops working. Nothing announces this. The link still exists on the other site and now points at nothing, which is worse than not existing.
Where they usually sit. On old addresses.
Links accumulate over time, so the oldest addresses often carry the most. Those are also the addresses most likely to be forgotten, which is a bad combination and the reason block two insists on completeness.
What to do with the list. Treat it as the priority.
Every address on it maps by hand, gets tested individually before launch and gets checked again afterwards. Nothing else in the audit earns that treatment.
Record What Is Already Broken
Almost every site has existing problems. Recording them before the migration is what stops all of them being attributed to it afterwards.
What tends to be there already. Ordinary accumulation.
Pages that were already slow, addresses that were already returning errors, duplicated content nobody noticed and sections that stopped performing months earlier for unrelated reasons.
What happens without the record. Everything becomes the migration's fault.
After a launch, any problem anybody finds is assumed to have been introduced by it. That is a reasonable assumption and frequently wrong. Without a record there is no way to demonstrate otherwise.
Why this protects the build team too. It cuts both ways.
A developer who inherits a site with existing faults, holding no record of them, will be blamed for all of them. Establishing the starting condition is as much in their interest as anybody's.
What to do with what you find. Fix it or note it.
Some faults are worth correcting before the move so they are not carried across. Others are worth recording and leaving, so that the migration is not held responsible for them.
Where The Data Comes From
Four kinds of source contribute to an audit. None is complete on its own. That last point is the one worth taking away.
A crawl of the site. What exists and what is on it.
The most complete picture of addresses and content, produced by walking the site as anything else would. It finds what is reachable and it will not find addresses nothing links to.
The search engine's own reporting. What it has actually seen.
Which addresses are known, which queries produce impressions and which pages receive clicks. This finds addresses a crawl misses, including ones nobody links to internally.
Analytics. What people actually did.
Which pages received visits and what happened next. This is the only source that connects addresses to commercial outcomes.
A link source. Who points at what.
Which addresses have external links, per block six. This is the source most likely to be incomplete, because no provider sees every link.
Why more than one is needed. Each has blind spots.
An address with no internal links and no traffic can still carry external links. The union of the sources is the inventory. Any single one of them is a partial view presented confidently. The mechanism for the search and analytics sources is covered in our guides for each of those platforms.
Keep The Old Site Somewhere
A copy or archive of the old site is worth having. It is trivial to arrange before launch and frequently impossible afterwards.
What it gives you. A reference.
The ability to look at what a page actually said, rather than at a record describing it. When somebody asks whether a section was shortened, this settles it in seconds.
What it is not. A rollback.
A copy is for reference rather than for restoring. Returning a site to a previous state after addresses have moved is a second migration, covered on the planning page.
Why it becomes impossible. The environment goes.
Hosting gets cancelled, the platform subscription lapses and access is lost with the handover. None of that is unreasonable and all of it happens quickly once the new site is live.
What to keep. The content, at minimum.
Even a stored copy of the pages is enormously more useful than nothing, at almost no cost to produce. The planning sequence is in how to plan a site migration, what to watch afterwards is in monitoring after migration. The full series is on the website migration guide.
Once it is gone,
the record
is gone.
Every address recorded rather than the pages somebody remembers, the benchmark taken page by page because totals conceal a collapsed section, the linked addresses listed separately because nothing else is irreplaceable, existing faults noted so the migration is not blamed for them, with a copy of the old site kept while that is still possible.
What a migration engagement covers:
A migration is a one-off project rather than a monthly service, so it is scoped and quoted against the size of the site and how much is changing at once.
Every guide.
One practice.
What a migration is, the types, when to do it, planning, the checklist, auditing, mapping, redirects, internal links, tags, downtime, crawling, structured data, performance, notifying, monitoring, recovery and the mistakes.