Open any SEO tool, type in a competitor’s domain, and you’ll be shown a monthly traffic figure with a comma in it. It looks like a measurement. It isn’t one, and building a content strategy on it is the most common expensive mistake in this whole discipline.
Competitor content analysis in 2026 is a ranking exercise, not a measuring exercise. You can reliably learn the order — who’s ahead, which of their pages does more than which other page, what’s growing and what’s decaying. You cannot learn the magnitude. Every useful conclusion in this article follows from that distinction.
The number at the centre of every competitor tool is off by about half
The most credible source here is Ahrefs’ own research, because it’s a tool vendor publishing unflattering findings about its own product.
Ahrefs compared US organic traffic estimates against actual Google Search Console data for 1,635 random websites. The median deviation was 49.52% — meaning most of the time the estimate misreports a site’s traffic by up to half its value. The spread is wider than that median suggests: for some sites the estimate is off by under 5%, for others by more than 1,000%. In the same study Ahrefs calculated Semrush’s median deviation at 68.36% (Ahrefs).
Independent work lands in the same territory. One analysis of 184 websites reported an average error rate of 61.58% for Semrush — overestimating in 112 of the 184 cases — and 56.95% for Similarweb, with one site estimated at 130,000 against an actual 50,000 (Collaborator). Practitioners with access to both tools and client analytics report the same pattern: a site doing 3,000 visits a month estimated at 15,000 (Topicfinder).
Small sites are worst affected, which matters because your competitors are probably small sites.
Why it’s wrong, and why it got worse
Every major platform builds the estimate the same way: crawl as many keywords as possible, attach an estimated search volume to each, check where the site ranks, predict a click-through rate for that position, then multiply and sum.
Each step introduces error and the errors compound. No tool knows every keyword a site ranks for. Search volume is itself an approximation. Rankings move daily.
But the fourth step is the one that broke recently. The CTR model assumes a search results page that increasingly doesn’t exist. AI Overviews and AI-generated answers sit above the organic results and absorb clicks that the model still attributes to position one. Ahrefs’ own 2026 measurement put the CTR drop at position one at around 58% where an AI Overview appears.
The consequence is specific and worth stating: these estimates are not merely imprecise, they’re systematically biased upward on exactly the keywords where AI answers appear — which tend to be the informational keywords your competitor’s blog is built on. A competitor whose content looks like it’s holding steady may be losing clicks the tool can’t see. The mechanics behind that decoupling are in writing SEO content with AI without sounding robotic.
Use ordering, not magnitude
Here’s the rule that rescues the whole exercise, and it comes from the same Ahrefs study.
Consistency survives even when accuracy doesn’t. If site A genuinely gets more traffic than site B according to Search Console, that ordering usually holds in the tool’s estimates too — regardless of how wrong both numbers are. The tool is a bad thermometer and a decent ranking device.
So convert every question into a comparative one:
- Not “how much traffic does their guide get” but “does their guide out-perform their case studies.”
- Not “they get 40,000 visits” but “they’re roughly four times us on this topic cluster.”
- Not “this page is worth $X” but “this page has been climbing for six months while that one has been decaying.”
And use one tool consistently rather than averaging three. Mixing sources destroys the one property you can actually rely on. This is the same discipline that governs sentiment scores in social media listening: the absolute figure is noise, the direction and the ordering are the signal.
What you can actually know
| Knowable exactly | Knowable as an ordering | Not knowable |
|---|---|---|
| What they published, and when | Which pages perform better than which | Actual traffic numbers |
| Page structure, depth, format | Which topics they’re winning | Revenue, conversion rate, margin |
| Internal linking and site architecture | What’s growing vs decaying | Email list size and quality |
| Publishing cadence | Roughly how far ahead they are | Paid spend behind a page |
| What they claim and cite | Which competitor leads a cluster | Whether any of it is profitable |
| Whether they update old posts | — | Why a page actually worked |
The right-hand column is the one people forget. A competitor’s top-ranking page might be a loss leader, a legacy accident, or a piece that brings in traffic that never buys anything. Ranking is not the same as working, and you cannot see the difference from the outside.
The second track nobody runs
Traditional competitor analysis answers “which of their pages rank.” In 2026 that’s half the job, because ranking and being cited by AI assistants have separated — a substantial share of pages cited in AI answers don’t rank in the top ten at all.
So there’s a second question with almost no tooling behind it: which of their pages do assistants actually cite?
Run it manually, monthly, in about an hour:
- Write fifteen questions a real buyer in your category would ask.
- Ask them across two or three assistants, in fresh sessions with no personalisation.
- Log which competitors get named, and which specific URLs get cited.
- Look for the pattern in what gets cited — usually it’s pages with original data, clear structure, specific numbers, and direct answers rather than long narrative build-ups.
What you find is frequently not their best-ranking content. It’s often a comparison table, a pricing breakdown, or a documentation page nobody optimised. That’s the actionable finding, and it’s invisible to every rank tracker you’re paying for.
Where AI genuinely helps
| Task | Why it works |
|---|---|
| Clustering their catalogue into themes | 200 URLs into 12 topic groups in minutes — tedious, mechanical, reliable |
| Extracting argument structure | What claim does each top page make, what evidence does it use, what does it leave out |
| Mining reactions | Comments, YouTube replies and reviews on their content — real language, nothing invented |
| Gap-finding | Comparing their coverage map against yours and naming what neither of you covers |
| Format analysis | Word count, heading depth, table and image use across their winners |
What AI cannot do here is tell you why something worked. It will happily produce a confident causal story — “this ranks because of the FAQ schema” — assembled from correlation and pattern-matching. That’s the same fabrication risk catalogued in fact-checking AI-generated content, applied to strategy instead of facts, and it’s more dangerous here because nobody checks a strategic claim against a source.
Treat every causal explanation as a hypothesis to test on one page, not a finding to roll out across thirty.
A workflow that produces decisions
- Pick three competitors, not ten. One ahead of you, one level, one behind. The one behind often tells you most, because you can see what didn’t work.
- Export their content inventory — URLs, titles, publish dates, estimated position — from one tool only.
- Cluster it with AI into topic groups and formats.
- Rank the clusters against each other using relative estimates, never absolute numbers.
- Read the top three pages in each winning cluster yourself. Not a summary — actually read them. This is where the insight lives and where AI adds least.
- Run the citation track described above.
- Write down what you’ll do differently and what would prove you wrong.
Step five is the one people skip and the one that pays. A model can tell you a page is 2,400 words with eight headings and a comparison table. It can’t tell you that the writer clearly ran the test themselves, and that’s why people link to it. The information-gain argument runs through using AI for market research as well: mining real text you can access beats inventing analysis of text you can’t.
The trap at the end of all this
The most common outcome of competitor analysis is deciding to publish what the leader publishes. That’s usually the wrong conclusion for two reasons.
You’d be copying their constraints, not their strategy. Their format may reflect a team of six, a design budget, or a decade-old domain. Reproducing the output without the underlying advantage produces a worse version of something that already exists.
The window has moved. Their winning page was published against a search landscape that no longer exists. What earned position one in 2023 was written for a results page without an AI answer sitting on top of it.
The useful output of competitor analysis is almost never “do what they did.” It’s “here’s the thing they can’t do, or won’t.” That’s the gap worth building a calendar around — the planning side of which is in building an AI-driven content calendar, with the production economics in turning one piece of content into ten.
If your content feeds a store, the structured-data angle in AI ecommerce tools that increase sales matters more than any competitor’s blog post, and for the wider stack, the best AI tools for small business owners is the map.
Frequently asked questions
How accurate are competitor traffic estimates?
Not very. Ahrefs’ own study of 1,635 websites found a median deviation of 49.52% from actual Search Console data, and calculated Semrush’s at 68.36%. Independent analyses report similar or larger average error rates, with small sites affected worst. Treat these figures as relative indicators rather than measurements.
If the numbers are wrong, what are the tools good for?
Ordering. Ahrefs’ research found that consistency holds even where accuracy doesn’t — if one site genuinely gets more traffic than another, that ordering usually survives in the estimates. So comparative questions work: which competitor leads a cluster, which of their pages outperforms which, what’s growing and what’s decaying. Absolute questions don’t.
Should I cross-check estimates across several tools?
For a single site, checking two or three gives you a sense of the range. But for ongoing competitor tracking, pick one tool and stay with it — mixing sources destroys the internal consistency that makes the ordering reliable in the first place.
Why have traffic estimates become less reliable recently?
Because the click-through-rate step in the estimation chain assumes a search results page without AI answers on it. Where an AI Overview appears, measured click-through at position one has dropped sharply, so estimates are biased upward on precisely the informational keywords most blogs are built on.
What can AI actually do in competitor content analysis?
Clustering a large URL inventory into themes, extracting the claim-and-evidence structure of top pages, mining comments and reviews on their content, spotting coverage gaps, and summarising format patterns. What it can’t do is explain why a page succeeded — it will produce a confident causal story from correlation, which should be treated as a hypothesis to test, not a finding.
How do I find out which competitor pages AI assistants cite?
Manually, for now. Ask fifteen realistic buyer questions across two or three assistants in fresh sessions and log which competitors and which specific URLs get cited. The cited pages are frequently not the best-ranking ones — comparison tables, pricing pages and documentation show up often — and no rank tracker will surface this for you.