What Is Programmatic SEO and When Does It Work?
Generating pages from a structure and a data source rather than writing each one by hand. The same method produces perfectly legitimate pages and the thing published policy calls scaled content abuse, with nothing separating the two except whether each page was worth making. We use this method, so the distinction matters to us as much as to you.
What It Actually Is
A page structure, a source of data and a process that combines them. Instead of somebody writing each page individually, pages are produced from a template filled with information that differs per item.
Why calling it a method matters. It has no opinion.
A method is neutral about what it produces. The same one builds a set of pages nobody should have made and a set that is exactly what a searcher wanted. Nothing in the machinery distinguishes them.
What people mean when they call it a strategy. Usually a volume target.
Described as a strategy, it becomes a plan to produce a quantity of pages. That framing supplies the number and leaves out the judgement, which is where every failure in this subject begins.
Where it came from. Necessity rather than cleverness.
Sites with thousands of genuinely distinct items could never have been written by hand. The method exists because handwriting a page for every property, product or route is not possible rather than because somebody found a shortcut.
Why the reputation is mixed. Both versions are common.
Plenty of the internet's most useful pages were produced this way. So were a very large number of pages that should not exist. The reputation reflects the distribution rather than the technique.
Where It Is Entirely Legitimate
Some subjects genuinely have one page per item. Writing those by hand would be arbitrary rather than better. That is the case for the method and it is a strong one.
What those subjects have in common. Real distinct facts per item.
A property has its own address, price, size, condition and photographs. A route has its own distance, duration and stops. A product has its own specification. Those differences exist before anybody writes anything.
Why hand writing would be worse. Inconsistency, not craft.
Thousands of pages written individually would vary in what they covered, so a reader comparing two of them could not. A consistent structure is genuinely better for the person using it.
The everyday examples. Familiar and uncontroversial.
Property listings, product pages, timetables, comparison pages between two named things, plus reference material where each entry has its own data. Nobody objects to any of that and nobody should.
What makes them work. The page answers a real question.
Somebody searching that specific item is glad the page exists, because it tells them something they wanted to know. That is the whole of the legitimacy and it is the test in block four.
And Where It Becomes Something Else
The same method producing pages nobody asked for, differing only by a substituted word, is what published policy describes as scaled content abuse. The machinery is identical in both cases.
What the substituted word version looks like. Familiar to everybody.
A page for every town, every trade, every combination of the two, where the template is fixed and one word changes. The output is one page and a large number of copies with different labels.
Why the method being identical matters. You cannot judge by the process.
Nobody can look at how pages were produced and conclude anything. The same script, the same template and the same data source produce both outcomes, so the assessment has to be of the pages rather than of the production.
What published policy actually says. Purpose over authorship.
Scaled content abuse describes producing many pages primarily to manipulate rankings rather than to help people. Whether a person or a machine produced them is explicitly not the point. Our risky practices material carries the attributed account.
The uncomfortable implication. Good intentions do not settle it.
A business producing town pages because everybody does and because it seemed like sensible marketing is described by that policy just as much as somebody doing it cynically. Intent to manipulate is not required.
The Test Is Whether The Page Was Worth Making
Would somebody searching that specific thing be glad this page exists. That is the test. A page passing it is legitimate however it was produced. One failing it is not, however carefully it was written.
Why one criterion rather than a checklist. Checklists get gamed.
Any list of properties becomes a set of boxes to fill. Pages get built to satisfy the list rather than the reader. A single question about the person arriving cannot be satisfied by adding elements.
What passing looks like. Specific relief.
Somebody lands on the page and finds the thing they were trying to establish, which they could not have got from a more general page. That is a recognisable feeling and it is worth imagining without flattering yourself.
What failing looks like. A general page with a name inserted.
Somebody lands and realises they could have read any version of this. The specific thing they searched appears in the heading and nowhere in the substance.
Why the test cuts both ways. It permits and forbids.
It licenses generating thousands of pages where each answers something real, which some cautious businesses avoid unnecessarily. And it forbids forty pages where the town is the only variable, which many businesses are currently paying for.
Where the same test appears elsewhere. Two other angles.
From the strategy side in why topical authority outperforms keywords. From the policy side, in our risky practices material. All three reach this question.
We Use This Method
This agency builds location and sector pages this way. That is our core production method, which means we have a direct interest in where this line falls and you should read the rest of this page knowing it.
Why we are saying so. The alternative is worse.
An agency explaining the risks of a method while quietly using it is asking to be found out. That finding would undermine everything else on this site. Saying it first costs less and is also simply correct.
What we have actually done about it. Measured our own.
We have compared the prose across our own generated sets to establish how much genuinely differs page to page, then reworked the ones that did not hold up. That is the check described in block seven.
What we will not publish. The figures.
Those came from client specific work. A percentage without its context is misleading rather than informative. The useful part is that the check exists and that anybody can run it.
What this does not mean. That our pages are all fine.
It means we look and act on what we find. Any business publishing at volume, including this one, is one careless decision from the wrong side of block three, which is why the check is periodic rather than a single event.
What Makes The Difference In Practice
Three properties. The useful thing about all three is that somebody can verify them rather than assert them.
Genuine data behind each page. Rather than a substituted name.
Something true of that item and not of the others, drawn from something real. Where the only per-page variable is a name, there is no data behind the set and the pages have nothing to differ about.
Something on each page that exists nowhere else. The strictest of the three.
Not a rearrangement of shared material. Information a reader could not obtain from any other page in the set, which is what makes a page worth having rather than worth skipping.
A structure worth reading by the person it is for. The whole page, not the top of it.
Somebody who reaches the bottom should have learned something. Pages where the distinct material appears in the first paragraph and the rest is shared boilerplate fail this while passing a casual glance.
Why checkable matters. It survives a change of staff.
Aspirations about quality do not outlive the person who held them. A property somebody can test is still a property in two years when different people are producing the pages.
How to apply them. Before the set is built.
All three are decisions about whether a set should exist, so they belong at the planning stage rather than as a review afterwards.
Measure Your Own Overlap
Compare the actual prose across a set of generated pages and establish how much of each is genuinely different. It is a real check and anybody producing these pages can run it. Almost nobody ever has.
What you are measuring. Substance, not templates.
Not the shared navigation, headers or footers, which are identical everywhere and should be. The body content, which is what a reader came for and what the set is claiming to differentiate.
Why the result is usually a surprise. Nobody has looked.
Businesses producing these sets have read one or two pages and assumed the rest resemble them. Comparing forty at once produces a different impression from reading two. It is frequently uncomfortable.
What to do with a bad result. Not necessarily delete.
High overlap means either the pages need genuine per-item material adding, else the set was too fine-grained and several should be one page. Consolidation is usually the better answer.
Why we recommend a check that could fail us. The alternative is advice we ignore.
Recommending a test we had not run on our own work would be advice we did not take. That is the reason block five exists in the form it does.
Why no threshold appears here. There is not one.
No published figure defines acceptable overlap and we are not inventing one. The finding is directional: if most of each page is shared, you know what you have.
Where It Fails Commercially Rather Than Technically
A great deal of programmatic content is not risky. It is simply useless. Thin pages nobody reads waste effort even where nothing penalises them. That argument is stronger than the risk one, because it does not depend on being caught.
What the risk framing invites. A calculation.
Presented as risk, this becomes a question about probability. Businesses reasonably conclude that plenty of sites do it without consequence. That conclusion is frequently correct and entirely beside the point.
What the useless framing removes. The other side of the trade.
If the pages produce nothing, the calculation has no upside to weigh against the risk. You are not gambling. You are spending on something with no return and a small chance of an additional problem.
What the waste actually consists of. More than production cost.
The building of them, the maintaining of them, the attention they consume that would otherwise reach pages that earn, plus the impression they contribute to what the site collectively demonstrates.
The version worth watching for. A set that once worked.
Sets built years ago that performed then and produce nothing now. They are rarely reviewed because they are not new. They are usually the largest single opportunity to improve a site.
How to establish which you have. Look at the set, not the total.
Whether the pages in a generated set actually receive anything, examined as a set rather than absorbed into a site total that hides them.
Data Is The Constraint
The method only works where genuine distinct data exists per page. Inventing the distinctions is where it goes wrong. A business without the data should not use the method at all.
What having the data means. It exists before the pages do.
Real facts, already recorded somewhere, that differ per item. If somebody has to generate the differences in order to justify the pages, the data does not exist and the set should not be built.
The commonest way businesses fool themselves. Manufacturing local detail.
Writing a paragraph about each town, produced by somebody who has never been there, because the set needs something distinct in it. That is inventing distinctions rather than having them.
What to do without the data. Build fewer pages properly.
A smaller number of pages, written individually, covering the subject rather than the permutations. That is the topical approach and it is a better use of the same budget.
The question before commissioning any set. Ask it plainly.
What do we genuinely know about each of these that we do not know about the others, then where is it written down. An answer naming a real source means proceed. Hesitation means the set should not exist.
Why we would say no. It would not work.
A set built on invented distinctions produces the useless outcome in block eight. Structuring what you do have properly is covered in what is information architecture in SEO. The full series is on the advanced SEO guide.
Was each page
worth
making?
We will measure the prose overlap across your generated set and tell you what we find, including where consolidation would serve you better than more pages. If you have nothing genuinely distinct per item, we will say the set should not be built rather than building it.
What the national tier covers:
The national tier is £1,550 a month. It suits businesses competing beyond one town. We will say so if a lower tier would serve you better.
Every guide.
One practice.
Semantic search, machine learning, topical authority, programmatic production, rendering, log files, crawl allocation, information architecture, content pruning and how strategy changes with scale.