How AI Crawlers Read Your Website: 167-Site Study
AI bots stopped browsing. They started asking.
I have spent the last year watching how AI systems interact with real websites through LightSite. The clearest change is simple: AI agents are becoming more direct.
They still crawl websites, but increasingly they do not behave like humans moving through navigation. They look for a specific fact, answer, product, or business detail and try to retrieve it with as little work as possible.
For marketers, that changes what “AI-ready” actually means. This study covers 167 websites observed between October 1, 2025 and July 31, 2026.
AI activity increased, but retrieval became more efficient
Across our comparable request data from March through July 2026, AI agent requests increased from 4,529,809 in March to 7,001,351 in July.
The more interesting change happened underneath that growth. Between May and June:
- One-shot, whole-profile reads increased from 250,826 to 690,111, up 168%.
- Smaller, granular skill calls fell from 160,360 to 40,965, down roughly 70%.
- Websites receiving targeted skill calls increased from 78 in March to 160 in June.
In other words, AI systems were reaching more websites while often using fewer individual retrieval steps.
That is the important signal. I do not think marketers should optimize for more crawler hits. I think we should optimize for how easily an AI system can get the correct answer. Request counts alone are a weak metric — AI bot behaviour tells you more than server logs do.
Different AI crawlers behave differently
Treating all AI bot traffic as one metric hides another important pattern. Since June, we have seen clear differences between engines.
| AI system | Bulk reads | Targeted skill calls |
|---|---|---|
| ChatGPT | 581 | 120,594 |
| Gemini | 210,752 | 40 |
| ByteDance AI | 980,085 | 0 |
| Claude | 170,134 | 1,584 |
Some systems behave mostly like broad readers. Others make much more targeted requests.
I would not claim these requests reveal exactly what happens inside each AI platform. We cannot see that. What we can observe is the retrieval behavior.
For marketers, that means “AI traffic” should not be treated as one audience. GPTBot, ClaudeBot, PerplexityBot and Google’s crawlers each arrive with different habits, and AI bot analytics is where those differences become visible. GA4 will not show you any of it.
AI agents disproportionately ask questions
The strongest targeted behavior we observed was question answering. Across the measured skill traffic, qa.search generated 520,882 requests — far more than individual categories such as product search, FAQs, testimonials, or traditional page browsing.
This matches what we saw when we first introduced LLM skills and agent search. AI assistants are built around questions.
A user asks which CRM suits a small business, whether a product works for sensitive skin, how two platforms compare, or what something costs. The agent wants the answer quickly.
This is why I believe marketers should stop thinking only in pages. A good page still matters, but the facts inside it need to work independently: product information, comparisons, FAQs, pricing, proof, company identity, and direct answers.
Business facts reached almost every site
Another pattern was breadth. Business identity, products, FAQs, categories, manifests, and testimonials were requested across a wide range of websites. Traditional page browsing was much narrower.
This reinforces something we have seen in our earlier structured-data research: machine-readable structure does not guarantee an AI recommendation, but it reduces the amount of guessing required.
For marketers, the practical question is simple: can an AI system clearly determine who you are, what you sell, and why someone should trust you?
If the answer requires five pages and several assumptions, the site is creating unnecessary friction.
A newer signal: ARD discovery
We saw another interesting signal after launching LightSite’s Agent-Readable Directory (ARD) on July 12, 2026. ARD gives an agent a catalogue of what information a website can provide. The same idea sits behind Google’s Agentic Resource Discovery specification, published in June 2026 as an open way for agents to find and verify the tools, skills, and resources a site exposes.
In its early observation period, we recorded 179 ARD requests across 51 websites, with every recorded request returning HTTP 200. ChatGPT-related agents, Meta AI, Claude, headless agents, and other automated clients all discovered it.
The volume is small, which is exactly why the behavior interests me. Agents generally appear to check the catalogue once or twice rather than crawl it continuously.
It looks more like discovery than browsing: find the available information, understand the structure, then move directly toward what is needed. That fits the broader pattern we observed throughout 2026.
What marketers should do
I would focus on four things:
- Write direct answers. Build content around real comparison, pricing, product, use-case, and objection questions.
- Expose clear business facts. Make products, services, company identity, FAQs, and proof easy for machines to retrieve.
- Reduce retrieval steps. Do not force an agent through unnecessary navigation or JavaScript to understand basic information.
- Measure behavior, not bot hits. Track what agents request, whether retrieval succeeds, and whether AI-referred humans follow.
You can test some of these foundations with our Generative Engine Optimization Checker.
Methodology
Study population: 167 websites that passed through the LightSite platform between October 1, 2025 and July 31, 2026.
Request trend window: The directly comparable monthly request series used in this article covers March 1 through July 31, 2026.
Source: First-party, server-side telemetry generated when AI-related crawlers and agents requested LightSite’s machine-readable website infrastructure.
ARD observation: ARD launched on July 12, 2026. The ARD data is a newer supporting observation and is not included in the March–July request trend.
Limitations: We can observe requests, retrieval paths, response status, and agent identification. We cannot see how an AI provider uses information internally, whether a request relates to model training, or whether one retrieval directly caused a citation or recommendation.
The data supports one conclusion clearly: AI systems are not simply crawling more. They are getting more deliberate about what they retrieve.