Comparison content was always off-limits for a brand to publish itself, until LLMs made it worth the risk. Here's what happened when I shipped a batch of best-of roundups with real editorial guardrails, and eight months of citation data to show for it.

The BOFU gap nobody wants to own

As a full-funnel content marketer, I've always known there's a real gap in bottom-of-funnel content: the comparison pages that are genuinely instrumental in a buyer's final decision. And for just as long, it's been off-limits for a brand to publish that content itself.

There's a simple reason. For "best X" and "X vs. Y" queries, Google typically ranks third-party review sites and Reddit threads well ahead of anything a vendor publishes about its own category. If you're unlikely to rank, and therefore unlikely to get traffic, there's little incentive to invest in the page. So for years, I stayed away from this format entirely.

What changed with LLMs

Asking an assistant "what's the best tool for X" is now one of the most common ways people comparison-shop. LLMs can summarize an entire category into one clean answer, exactly the kind of answer a buyer wants when they're choosing between options.

When I started scoping which prompts were worth tracking and targeting for GEO, BOFU comparison queries stood out immediately. They reveal buyer stage and intent about as clearly as a query can. And it was already visible in the wild that LLMs were citing and pulling from "best X" content directly in their answers. Of everything I could have tested first, this looked like the highest-leverage bet.

What I ran

I published a batch of just under fifty articles targeting "best X" style queries, each featuring our own product alongside category peers. Because this content sits closer to the edge of what a brand can credibly publish about itself, I built in real guardrails before any of it shipped:

  • No 1:1 comparisons. We leaned on round-up posts instead, a format that performs well in AI search, is industry-standard, and is what most competitors run anyway.
  • Neutral peer mentions. Copy focused on where we differentiate, not on calling out competitor limitations.
  • Legal review on every article before it published.
  • Ongoing performance monitoring, with SEO retiring any page that didn't hit its traffic/signup threshold.
  • A quarterly audit and refresh cycle, given how fast this space moves.

Then I tracked citations per page, week over week, for eight months.

The citation ramp

For the first few weeks the pages barely registered. The week the full batch shipped and got indexed, citations inflected almost straight up. Within a quarter, weekly citations were up ~20× off the starting baseline; at peak, ~450×.

Weekly citations, indexed to launch
Sampled across an eight-month run · launch week = 1×
0 100× 200× 300× 400× batch ships ~20× (1 quarter) ~450× start peak 8 mo
Anonymized and indexed to the launch week (= 1×); the shape reflects the actual program. Traffic and downstream metrics are omitted. Bars before the dashed line are the pre-launch baseline.

Which pages carried it

Citations weren't evenly spread. A handful of roundups did most of the work, and the pattern behind them was clear: the pages that performed best sat in categories where our product already had strong brand association. Where recognition was thinner, the same format did far less.

Page archetype Brand association Citation index
Best-of roundup, flagship categoryStrong100
Best-of roundup, adjacent categoryStrong64
Top-tools roundup, secondary categoryModerate33
Best-of roundup, secondary categoryModerate27
Category roundup, unfamiliar categoryWeak11

Relative citation index. Top page = 100. Archetypes, categories, and values are anonymized.

There was a second pattern worth calling out. Around the same window, I'd also published a set of Use Case pages to support a new product launch. The best-of listicles that sat in that same product category picked up an extra lift in citations that tracked closely with when those Use Case pages went live: two different content types, in the same topical neighborhood, reinforcing each other in front of the same model.

The caveats

  • Citations plateau, then dip. After the peak, weekly citations settled back ~10%. The timing lines up with a Google AI Overviews citation algorithm change around then, so I'd read most of the dip as a platform shift rather than the pages decaying on their own. Still, it's a good reminder that generative visibility moves with the engines, not just your content.
  • Distribution is lopsided. A few pages carried most of the citations; the long tail earned little. Publishing the batch didn't mean every page in it would win.
  • A citation isn't a conversion. Getting quoted is the top of a longer, murkier funnel. Read citation growth as a visibility signal, not revenue.
  • The format can be abused. Thin listicles still read as thin. Structure earned the citations; having your own product in the list just didn't disqualify them. The best performing articles are still within the category that your product actually performs well in, or branch into secondary categories.

What I learned from this experiment

  • Build in the guardrails first, not after. Neutral peer treatment and a legal review pass make this defensible before it's a problem.
  • Structure for extraction. One claim per heading, a direct answer near the top, comparison tables, clean markup.
  • Aim at live questions. Prioritize what people are actually asking assistants, not just high-volume keywords.
  • Ship in a batch so the engines pick up the set in one crawl.
  • Measure citations directly, per page, week over week. Retire what doesn't perform, and expect a plateau so you know when to refresh.
  • Attribution is muddy. While the pages captured citations, it's hard to attribute down-funnel performance, particularly new user signups. It may be that users searched the brand name and landed on the homepage for a conversion journey, but more needs to be tested there.
Get in touch ← All resources