Benchmarks vs Performance Data in AI Search
Every AI search deck now opens with an industry number. Share of citations by platform. Average visibility by vertical. They are useful for orientation and almost useless for deciding what to do on Monday.
A benchmark tells you where the market sits. Performance data tells you what your site did. Only the second one can be acted on.
What benchmarks are for
Benchmarks set expectations and win arguments about budget. Two examples we use ourselves:
- In a scan of 2,870 websites, 27% were blocking at least one major LLM crawler — usually at the CDN or WAF layer, not in robots.txt.
- Across the sites we measure, roughly 12% of pages receive about half of all observed AI attention.
Both are good slides. Neither tells you whether your crawlers are blocked, or which of your pages are carrying the attention.
What performance data looks like
Performance data is first-party, from your own site, over a defined window. It answers narrower questions with numbers you can move:
- Two of our measured properties received 95 and 23 human visits per 1,000 AI crawls. Same category, same crawl volume band, four times the return. The gap was in the pages being crawled, not in crawl volume.
- On one enterprise property, 7,707 AI-referred visits split 83% to support and documentation pages and 6% to the blog. Content investment was going almost entirely to the 6%.
That second finding changed a roadmap. No benchmark would have produced it.
How to use both without confusing them
- Use benchmarks to frame the question. "A quarter of sites block a crawler" is a reason to test your own.
- Use your own data to answer it. Run the check, get your number, and treat it as the baseline.
- Compare yourself to yourself. The only reliable comparison is the same site across two windows with one deliberate change between them.
The failure mode is reporting benchmarks as if they were results. It feels like measurement and produces no decisions.
If you want the longer version of how first-party AI traffic gets read, we cover the measurement stack in AI search measurement and the behavioural signals in what AI bot traffic can actually tell you. To get your own baseline, run the free AI visibility checker.
FAQ
Are AI search benchmarks unreliable?
They are reliable as market context and unreliable as a target. Sampling, verticals and crawler mixes differ too much to treat an industry average as your goal.
What is the minimum performance data to start with?
Crawler access status, which pages get crawled, and human visits per 1,000 crawls. Those three cover access, attention and return.
How long should a measurement window be?
Thirty days is usually enough to see a behavioural change and short enough to keep the comparison clean.