What Is Thin Content in SEO?
Unlike everything else in this section, this is the risk almost every site actually carries. It is rarely deliberate, most sites have some and having it is not a scandal. It is also the most fixable thing here. The fix is usually the opposite of what people expect.
This Is The One You Probably Have
Every other page in this section describes something somebody chose to do. This one describes something that accumulates while nobody is watching, which is why it is both the commonest and the least alarming.
How it arrives. Ordinary decisions.
A page created for a campaign that ended. A location page added because a competitor had one. A category with two products in it. A post written to have something to post. None of that was careless at the time.
Why nobody notices. Nothing announces it.
Pages do not degrade visibly. A page that was reasonable when there were forty of them reads differently when there are four hundred. No individual page changed.
What makes this page different in tone. Nobody is at fault.
The rest of this section deals with manipulation. This deals with accumulation. The appropriate response is a tidy rather than an investigation.
Why it still matters. It affects everything else.
A site where most pages do little makes it harder for the pages that do something to be found, which is a structural problem rather than a punishment.
Thin Is Not About Length
Word count is not the measure of anything. A short page that answers a question completely is not thin. A long page that says nothing is. Length has been used as a proxy for quality for so long that this needs stating plainly.
Where the confusion came from. Length is measurable.
Usefulness is a judgement and word count is a number. Given a choice between a hard judgement and an easy number, an industry will reach for the number every time, then start treating it as the thing itself.
What that produced. Padding.
Pages inflated to reach a target, with a history section nobody wanted and three paragraphs restating the question. That makes a page longer and worse simultaneously, which is the exact opposite of the intention.
What we will not give you. A threshold.
No word count makes a page acceptable or unacceptable. Anybody offering you one has invented it. The number will be whatever their tool defaults to.
The replacement test. Did it finish the job.
Whether somebody arriving with a question leaves with an answer. A page can do that in two hundred words or need two thousand. The subject decides rather than the target.
What Thin Actually Means
A page that does not do enough for the person who arrived on it to justify its existence. That is a judgement about usefulness rather than a threshold. It is deliberately harder to apply than a number.
The question that operationalises it. Would you send somebody here.
If a customer asked you this, would you send them this page or would you explain it yourself. Anything you would not send is thin regardless of how it looks.
The commonest shape. A page that exists to be a page.
Created because the structure seemed to require it, else because a list of topics said so. It states what it is about and stops, because nobody ever had anything to say about it.
The second commonest. Written for a search rather than a reader.
Assembled from what other pages on the subject contain, with nothing added. Perfectly readable and completely unnecessary, since everything on it exists elsewhere already.
Why this is not about effort. Hard work can produce it.
A page somebody spent a day on can be thin if the subject did not warrant a page. That is uncomfortable and it is why consolidation is easier to justify than deletion, which block ten covers.
Duplicate Content
The widely feared duplicate content penalty is not a thing. Google's own guidance states that duplicate content on a site is not grounds for action on that site unless it appears that the intent of the duplication is to be deceptive and to manipulate search engine results.
Why duplication is normal. Most sites produce it without trying.
Printer-friendly versions, addresses with parameters attached, product variants, syndicated material, a description repeated across similar items. None of that is deceptive and none of it is being punished.
What actually goes wrong. Signals get split.
Where the same content sits at several addresses and they cannot all be identified as versions of each other, their properties cannot be consolidated. That dilutes ranking signals across the addresses rather than combining them.
The realistic worst case. The wrong version appears.
Google describes the practical downside as a less desirable version of the page being shown in results. That is an inconvenience rather than a sanction.
The other real cost. Attention.
On a large site, having many versions of the same thing consumes crawling that could have gone to pages that matter. Our technical material covers handling that properly.
What that adds up to. Tidy it, do not fear it.
Duplication is a housekeeping matter with a mechanical fix rather than a risk requiring a defence.
The Template Problem
Pages generated from a template with one variable swapped are duplication in substance even where every word differs slightly. Rewriting the sentences does not change what the page is.
Why the wording is not the test. Substance is.
Two pages can share no identical sentence and still contain the same information arranged the same way. A reader moving between them learns nothing new, which is the thing that matters rather than whether a checker reports a match.
What the vendor says about this directly. Worth quoting the position.
Google's guidance says that deciding to build a site whose purpose inherently involves content duplication is something to think twice about where the business model will rely on search traffic, unless a great deal of additional value can be added for users. That is a description of templated trade and town pages, from the vendor rather than from us.
Why we are saying so. We build these.
This agency produces trade and town pages at volume. The condition in that guidance is the whole job. Additional value has to actually be added. Where it cannot be, the page should not exist.
Where the related arguments sit. Two other angles.
The doorway version is in what is black hat SEO. Scaled content abuse with the overlap check is on the Google spam policy page.
Scraped Content
Republishing somebody else's material without adding anything. It is one of the named areas in the published spam policies. It runs in two directions depending on which end of it you are.
Doing it. A clear policy breach.
Taking content from elsewhere and publishing it, whether copied outright or lightly reworded, is described directly in the policies. Republishing with permission and attribution is a different thing, provided the page adds something of its own.
Having it done to you. Common and usually harmless.
Established sites get copied constantly. Being the original, established source generally means the copy is not competing with you. Most instances sit unnoticed on sites nobody visits.
What is worth doing about it. Ask, then escalate if it matters.
Asking for removal frequently works. Where it does not, copyright complaint routes and search engine removal processes exist. Both are formal requests rather than technical fixes.
When to bother. Rarely.
If a copy is outranking you for your own material, pursue it. If it is sitting somewhere nobody visits, the time is better spent on your own site.
Content Produced At Scale By Machine
The objection is not that a machine was involved. It is content produced in volume without judgement, without review and with nothing to add. Published policy makes exactly the same distinction.
What the policy actually says. Purpose, not authorship.
Scaled content abuse describes producing many pages primarily to manipulate rankings rather than to help people. Whether a person or a tool produced them is explicitly not the point.
Why that matters. It removes a defence and a fear at once.
Having a human write templated pages does not take them outside the policy. Equally, using a tool to help produce something genuinely useful does not put you inside it. The pages either do something for the reader or they do not.
What the problematic version looks like. Volume without a decision.
Hundreds of pages generated against a keyword list, published without anybody reading them, on subjects the business has no particular knowledge of. Nobody decided each page was worth existing.
The specific risk in this trade. Plausible and wrong.
Generated material about regulated, technical or safety-related subjects can read confidently and be incorrect. Publishing it under your own name means carrying the consequence of it.
AI Tools Support A Writer Rather Than Replace One
This is the position we hold across every part of this site. It is worth stating in full here, because this is where it is most likely to be tested.
What responsible use looks like. Judgement stays with a person.
Structuring an argument, drafting something a person then rewrites, tidying prose somebody produced, getting past a blank page, checking whether anything obvious has been missed. In each case somebody decides what is true and what is worth saying.
What irresponsible use looks like. Publishing output.
Generating pages and putting them live without anybody deciding they should exist or checking whether they are correct. The failure is the absence of a person rather than the presence of a tool.
The test that separates them. Who is accountable.
If somebody could defend every claim on the page, it was written properly however it was drafted. If nobody could, it should not be published regardless of who or what produced it.
Why we are explicit about it. The question gets asked.
Businesses want to know whether using these tools is dangerous. The answer is that it depends entirely on whether anybody is exercising judgement, which was also true of every previous shortcut this industry has been offered.
The Nine Hundred Page Problem
A business can accumulate a very large number of pages that each seemed reasonable when it was made. Collectively they read as low effort. The figure in the heading is illustrative rather than a measurement of anything.
How a site gets there. One decision at a time.
Nobody sets out to build an archive nobody defends. It arrives through years of campaigns, trades added, towns added, posts written to maintain a schedule and sections nobody removed when their reason ended.
Why the total is what matters. The picture is cumulative.
Each page can be defensible while the whole reads as volume for its own sake. That is a judgement about the site rather than about any page in it, which is why reviewing pages individually never finds the problem.
Why we can describe this. We have looked at our own.
An agency producing trade and town pages at volume across many sectors is exactly the kind of publisher this describes. We have examined our own work on that basis and reworked what did not hold up. We publish no counts or percentages, because those came from client specific work.
Why it is fixable. Nothing here is a sanction.
No policy has necessarily been breached. This is a site that has grown untidily. Untidiness is reversible in a way that manipulation is not.
How To Fix It
Improve, consolidate or remove. That order is deliberate. Consolidation is the one businesses reach for last while it is usually the right answer.
Improve. Where the subject deserved a page.
If the topic warrants one and the page simply did not deliver, finish the job. This is the best outcome and the least common, because most thin pages are thin for structural reasons rather than for lack of effort.
Consolidate. The preferred fix, also the most overlooked.
Six weak pages on adjacent subjects become one page that covers the ground properly, with the old addresses pointed at the new one. Nothing is lost, whatever the pages had earned is preserved and the result is genuinely better.
Remove. Legitimate, though less good.
Some pages should simply go. Removal is a valid answer and it discards whatever the page had accumulated, which is why it sits third.
What not to do. Block access to the duplicates.
The instinctive fix is to hide duplicate versions from crawling. Google's guidance is explicit that this is wrong: versions that cannot be crawled cannot have their signals consolidated, so blocking them prevents the very thing you wanted.
The rule for the addresses. Handle them properly.
Anything consolidated or removed leaves an address behind. Our site migrations material covers redirecting them. The full series is on the dangerous SEO practices guide.
Consolidate.
Almost never
add more.
A review of what has actually accumulated on your site, which pages are doing work, which could be combined and which should go. Including the overlap check on templated pages. We will tell you plainly if we find nothing worth changing.
What a content review covers:
A content review is a one-off piece of work rather than a monthly service, so it is scoped against the size of the site and how much has accumulated.
Every guide.
One practice.
Black hat, grey hat, link risk, thin content, the spam policies, negative SEO, manual actions, algorithmic penalties, how to check for one and how to recover from one.