We wanted to build a programmatic SEO program that could scale with AI, without giving up the quality of our best pages. So I ran the closest thing I've had to a controlled test of what that tradeoff actually looks like, and tracked it for a full year.
The setup
We'd identified a long list of topics, each mapped to a specific jobs-to-be-done. Rather than write all of them by hand, or hand the whole batch to AI, we split it by opportunity:
- We allocated the thirteen highest-potential topics by demand to our competent human writer. These writings became the basis of the program, and they set the voice, structure, and depth we wanted everything else to match.
- Everything else, well over 150 topics, was generated programmatically by our AI system, using the human-written pages as the pattern to follow.
- Before scaling past a first batch, our product marketing manager reviewed the initial 50 AI pages for accuracy and tone. That feedback went back into the generation process before the rest shipped.
The caveat here is that this isn't a clean AI-vs-human test. We deliberately put humans on the topics we already believed had the most upside, and let AI cover the rest. So any gap, or lack of one, between the two groups is partly a story about where we placed our bets, not purely about who wrote the page.
Then I tracked traffic, signup conversion, downstream product usage, and a revenue proxy, per page, for twelve months.
The cohort split
The human-written set carried the large majority of the volume, which is exactly what you'd expect from prioritizing by search opportunity: 13 pages, about 7% of the total page count, brought in 89% of total traffic and 93% of total signups. The AI-generated pages, more than ten times as many pages, made up the remaining 11% of traffic and 7% of signups.
Per-signup, the two cohorts diverge more than I expected, just not in one consistent direction. Indexed to the human cohort at 100:
- Signup: AI pages indexed around 61, notably behind human
- Product usage: AI pages indexed around 118, ahead of human
- Revenue: AI pages indexed around 125, ahead of human
Fewer visitors converted on the AI-generated pages. But the ones who did convert were more likely to actually use the product afterward, and looked more valuable on the revenue proxy. The AI cohort's weakness was getting the first conversion, not the quality of the people who converted.
Organic rankings metrics
Looking specifically at organic search performance on nonbranded queries, for pages mature enough to have settled in the index, the picture is far less balanced.
Indexed to the human-written baseline at 100:
- Clicks per page: AI pages indexed in the low single digits, functionally negligible next to the human set
- Click-through rate: AI pages indexed around 55, roughly half the human rate
- Ranking strength: AI pages indexed around 80, ranking meaningfully lower on average
- Share of pages with meaningful organic clicks: every human-written page cleared a basic visibility bar this quarter; only about half of the mature AI pages did
It would be remiss if I don't caveat that the human-written pieces are the ones I prioritize for internal linking and optimizations, because at the end of the day, I am looking to drive the most business impact :)
A handful of AI pages still beat our weakest human-written pages
The AI cohort's best individual page didn't rival our top human-written pages, but it did outperform three of our thirteen human-written pages outright. Nothing about how it was written separated it from the rest of the long tail; it landed in a topic with more real demand than we'd assumed when we chose the human set.
That's a smaller version of the same point: topic selection is doing a lot of the work in this traffic split, not some inherent ceiling on what an AI-generated page can achieve.
Both cohorts ramped and peaked on a similar timeline
Both cohorts ramped on a similar timeline; the AI pages, once past the QA gate on the first 50, climbed in step with the human set rather than trailing far behind. Both also peaked within a month of each other and pulled back afterward, the human set down about a third off its peak by the most recent month, the AI set down closer to 37%. Close enough in timing and shape that I'd read it as a platform-level shift, the same read I gave the dip in the listicles data, rather than either cohort's content decaying on its own.
Topic selection mattered more than who wrote the pages
We prioritized the highest-potential topics for humans and used AI for the rest, so the traffic split reflects that prioritization more than it reflects a gap in output quality. Once a topic had real demand and mapped to something the product could actually solve, it performed, whether a person or a model wrote the first draft.
I think that's the actual lesson, and it's easy to miss if you only look at the human-vs-AI split. The generation method decided how fast we could produce the page. It didn't decide whether the underlying premise, does this product solve this person's problem, was true.
What I learned from this experiment
- Intent matters more than who or what wrote the page. Job-to-be-done and product fit did more work than generation method. Test that first before scaling the wrong topics fast.
- The split was a resourcing decision, not a quality test. Be honest about that caveat before drawing conclusions from the comparison.
- Use your best content as the pattern, not just a template. The human-written set wasn't just quality control, it was the training signal for everything that came after.
- Build in a manual QA gate before scaling. Reviewing the first 50 AI pages and feeding that feedback back in mattered more than any prompt tweak done in isolation.
- Watch for outliers in the long tail. One AI page rivaling your best human pages is a signal about topic selection, not a fluke to explain away.