What Is an XML Sitemap?
The most over-rated file in technical SEO. It helps pages be found and does almost nothing else, which is worth establishing before anybody spends money on the assumption that it does considerably more.
It Is A Suggestion, Not An Instruction
A sitemap tells a search engine which pages exist. It does not compel anything to be crawled, does not compel anything to be stored and has no bearing at all on whether a page is shown.
What it is. A list of addresses.
A file naming the pages you want found, usually with a note of when each last changed. That is genuinely the whole of it.
What it asks for. Consideration.
It offers a list and nothing more. Whether the pages are fetched, in what order and whether any of them is retained are all decisions made elsewhere.
What people believe instead. That it guarantees something.
The commonest assumption is that a page in a sitemap will be indexed. That is not how it works and believing it produces the misdiagnosis block ten describes.
Why the belief persists. It sounds official.
A file submitted directly to a search engine feels like a formal declaration. It is closer to leaving a list on a desk than to filing a form.
Where it sits in the sequence. The found stage only.
In the three stages on crawling and indexing, this helps only with the first. Nothing it does touches storage or presentation.
What It Actually Helps With
Three situations. Outside them the difference is very small. Knowing whether you are in one of them determines whether any of this deserves your attention.
Large sites. Where discovery takes time.
A site with many thousands of pages benefits from a list, because reaching everything by following links takes considerably longer than reading an inventory.
New sites with few links. Where nothing points inward.
A site nobody links to yet is hard to discover. A list gives something to start from rather than waiting to be stumbled upon.
Poorly connected sites. Where pages are orphaned.
Pages reachable only through a search box, a form or a deep chain of clicks. A list surfaces them, although the better fix is linking to them properly.
Where it barely matters. A small well linked site.
A business site of modest size with sensible navigation is entirely discoverable without one. Having a sitemap is still reasonable and expecting it to change anything is not.
Why we say that plainly. It gets sold as a fix.
Sitemap work is easy to deliver and easy to invoice. On most of the sites we see it was already there, generated automatically, requiring no work at all.
The better version of the orphan fix. Internal links.
A page nothing links to has a structural problem that a list conceals rather than solves. Linking to it addresses discovery and gives visitors a route.
What Should Be In It
The pages you want found and indexed, nothing else. Stated that simply it sounds obvious. The number of sitemaps that fail it is remarkable.
The test for inclusion. Would you send somebody here.
If a page is one you would happily direct a customer to, it belongs. If not, its presence is asking for attention on something you do not want attended to.
What that usually means. Fewer pages than you have.
Service pages, location pages, guides, products and the main structural pages. Not every address the site can produce, which is frequently a much larger number.
What the date note is for. Signalling change.
A record of when each page last changed, which helps prioritise revisiting. Only useful where it is accurate, which is undermined by platforms updating it whether the content changed or not.
What to do about size. Split it.
Very large sites use multiple files with an index listing them. That is a mechanical matter your platform handles rather than a decision you need to take.
What has no effect. Priority values.
The file format allows a stated importance for each page. It is widely ignored, so setting it carefully achieves nothing measurable and is not worth anybody's afternoon.
What To Exclude
Four kinds of page should not be listed. Including any of them makes the site say two different things at once, which is block five.
Blocked pages. The clearest contradiction.
A page your own instructions ask crawlers not to fetch has no business being advertised as a page you want found. That is a direct conflict with what your robots file says.
Redirected pages. The stale ones.
An address that now sends visitors elsewhere is not a page. Listing it points at something that no longer exists in its own right.
Pages that have gone. The leftovers.
Removed pages still listed produce a file full of addresses that return nothing. Common after a rebuild, where the list was built once and never revised.
Pages marked not to be indexed. The subtle one.
A page carrying an instruction that it should not be listed, while appearing in a list of pages you want listed. Both instructions are yours and they disagree.
Why exclusion is harder than inclusion. It needs knowing the rest of the site.
Deciding what belongs is a judgement about the page. Deciding what to exclude requires knowing what your other instructions already say, which is why this half gets skipped.
What the symptom looks like. Reported inconsistencies.
Reporting flagging listed pages as excluded, blocked or redirected. That is the file disagreeing with the site rather than a fault in either.
Contradicting Yourself Is The Common Fault
A sitemap listing a page that is blocked, redirected or excluded elsewhere is the site saying two different things. That is the commonest real problem with this file. It is also one of three places the same failure appears.
What the contradiction does. Leaves the decision open.
Given two conflicting statements from the same site, something has to choose which to act on. You have handed away a decision you could have made yourself.
Why it is worse than being wrong. Unpredictability.
A clear instruction you disagree with produces a predictable outcome. Two instructions that conflict produce an outcome nobody can anticipate, which is harder to diagnose later.
Where else this appears. Two other pillars.
A canonical disagreeing with a redirect, plus international declarations that are not reciprocal. Both are the same failure in a different mechanism, which the canonical pillar names in the same terms deliberately.
Why it happens. Different people, different times.
The sitemap is generated by a platform, the blocking was added by a developer and the indexing instruction was set by somebody in a hurry. None of them was wrong alone.
What consistency requires. Reading all your instructions together.
Checking what every mechanism says about the same page, rather than each in isolation. That is the whole job and it is rarely anybody's responsibility.
Who Maintains It
Most platforms generate this automatically, which is usually the right answer. A hand-built file goes stale the moment somebody publishes a page.
Why automatic wins. It keeps up.
A generated file reflects the site as it currently is. Every page added appears, every page removed disappears, then nobody has to remember anything.
What automatic gets wrong. Sometimes the exclusions.
A platform listing everything it can find may include pages you would rather it did not. That is worth checking once. It is a smaller problem than a stale file.
Why hand-built fails. Nobody updates it.
A file assembled manually is accurate on the day it was made. Six months and thirty pages later it describes a site that no longer exists. Nobody notices, because nothing breaks.
When manual is defensible. Very rarely.
Where a platform genuinely cannot generate one, else where a specific subset needs listing separately. Both are unusual and both need an owner.
What to check on an automatic one. Once, then leave it.
Whether it lists roughly the number of pages you expect and whether anything obviously wrong is in it. Ten minutes at launch, rather than a recurring task.
Submitting It
Submitting tells a search engine where the file is. It does not cause anything to be crawled. Its real value is that it turns on reporting about the file.
What it does. Registers a location.
You state where the file lives so it can be found without being discovered. Worth doing once and largely uninteresting after that.
What it does not do. Prompt anything.
Nothing gets fetched because a file was submitted. Submission is not a request for attention, which people reasonably assume from the word.
Why it is still worth doing. The reporting.
Once registered, you can see how many listed pages are known and how many are excluded. That comparison is the genuinely useful part and it is how block five's contradictions surface.
What the report tells you. Two numbers and a reason.
How many pages were submitted, how many are stored and why the difference exists. That last part is the diagnostic value. Everything else is noise.
Where the mechanism sits. The other cluster.
Where to submit and how the reports behave belongs in our Google Search Console material rather than here.
Image, Video And News Sitemaps
Three specialist variants exist. Almost no business we work with needs any of them, which is why they occupy one short block rather than three pages.
Image sitemaps. For image-led businesses.
A separate listing of images, useful where images are the product rather than illustration. Stock libraries and galleries, rather than a business with photographs on its pages.
Video sitemaps. For video hosted on your own site.
Describing where videos live and what they contain. Largely unnecessary where video sits on a platform, which already describes it thoroughly.
News sitemaps. For genuine publishers.
A listing of recent articles for sites publishing news frequently. This means an actual news operation rather than a business with a blog.
Who needs one. A small set.
Publishers, media libraries and businesses whose media is the offering. If you are unsure whether you qualify, you do not.
Why they get recommended anyway. They sound thorough.
Adding three extra files reads as diligence on a proposal. On a business site with ordinary photographs it produces three more things to maintain and no outcome.
Sitemaps During A Migration
During a move, this file stops being decorative. Every address changes, so the sitemap becomes one of the few things that can state the new arrangement directly.
Why it matters more. Discovery is the problem.
After a restructure, the new addresses are unknown and the old ones are redirecting. Anything that shortens discovery genuinely helps at that moment, unlike on a settled site.
The commonest failure. The old file survives.
A sitemap listing addresses that all now redirect, left in place because nobody thought about it. That is block five's contradiction applied to every page at once.
What should happen. A new file, promptly.
The new addresses listed as soon as the site is live, so what you want found is stated rather than inferred from redirects.
The argument about the old one. Handled elsewhere.
Whether to keep the previous file available for a period, so redirects are discovered faster, is a genuine question. Our site migrations material owns that argument.
What to check afterwards. The submitted count.
Whether the number of pages submitted matches what you expect the new site to contain. A large discrepancy is the first sign something was left behind.
It Will Not Fix Anything
A page not being indexed is rarely solved by adding it to a sitemap. The reason is almost always elsewhere. This is a common and expensive misdiagnosis.
Why the assumption is made. It feels causal.
The page is not indexed, the sitemap is the thing that lists pages, therefore listing it should help. The reasoning is tidy and the premise is wrong.
What the actual reasons usually are. Three.
The page is thin or near-duplicate and was judged not worth storing. Something on your site is blocking or excluding it. Or it is stored and simply never shown for anything.
How to tell which. Ask what stage failed.
Whether the page has been fetched, whether it was retained and whether it appears for anything. Those three answers point at three different fixes and none of them is a sitemap.
What the misdiagnosis costs. Time and the real problem.
Weeks spent waiting for a listing to take effect, while the actual cause remains. That delay is the expensive part rather than the work itself.
When it genuinely is the answer. Rarely, though identifiably.
A page nothing links to, on a site with no other route to it. That is a discovery failure and a list does address it, though linking to the page addresses it better.
What to do instead. Diagnose first.
The three-stage sequence on crawling and indexing identifies which stage failed. The full series is on the technical SEO guide.
Adding it to the
sitemap will not
fix it.
A page not being indexed is thin, blocked, excluded, else stored and never shown. Those are four different problems and a list addresses none of them. We establish which stage actually failed rather than waiting weeks for a listing to take effect.
What we check:
On most sites this file was already there and needed no work at all.
Every guide.
One practice.
How search finds and stores a site, what your own files are telling it, addresses and duplication, status codes, mobile, hosting and what a real audit contains.