August 5, 2026

Crawl-to-Refer Ratio: AI Bots Take 4,580 Pages for Every Visit (2026)

Q2 2026 AI agent requests hit 17.7 billion, up 45% QoQ. Here is what the crawl-to-refer ratio reveals and how to audit crawler access.

Anthropic's crawler fetched roughly 4,580 pages for every one visitor it sent back in June 2026. Google's crawler fetched five. That gap has a name: the crawl-to-refer ratio, and it is the sharpest number in AI search for showing what your content costs you against what it returns.

The volume behind the ratio jumped again last quarter. DataDome logged 17.7 billion AI agent requests across April, May and June 2026, up 45% from 12.2 billion in Q1. Your Google Analytics dashboard recorded none of it.

Here is what the Q2 numbers show, why the crawl-to-refer ratio splits so hard by platform, which crawler decides whether ChatGPT cites you, and what to change before Cloudflare's September 15 default block lands.

What Is the Crawl-to-Refer Ratio, and Why Did It Spike in 2026?

The crawl-to-refer ratio divides the pages an AI crawler fetches from your site by the referral visits its parent platform sends back. A ratio of 5:1 means the bot took five pages and returned one human. A ratio of 4,580:1 means it took 4,580 pages and returned one.

Cloudflare Radar data from June 2026 puts Anthropic near 4,580:1, OpenAI at 848:1, Perplexity at 186:1, and Google at 5:1. The spread is not a rounding artifact. It maps to product design.

Cloudflare's May 2026 breakdown shows why the ratio climbed: search-purpose crawling, the kind that can produce a citation with a link, made up under 10% of AI crawler requests. The remaining 90% went to training and answer generation, where the model responds in place and the user never leaves the chat window.

Track this metric for your brand → nobori.ai

How Much AI Agent Traffic Actually Hit the Web in Q2 2026?

DataDome's Q2 2026 AI Traffic Report, published July 16, counted 17.7 billion AI agent requests across its network in April, May and June. Q1 came in at 12.2 billion. That is 45% quarter-over-quarter growth in three months.

The monthly curve steepened rather than flattened:

  • April 2026: 4.77 billion requests
  • May 2026: 6.29 billion requests, up 32% in a single month
  • June 2026: 6.60 billion requests, the highest month on record

Fastly measured the same shift from a different vantage point. AI requests on its platform grew about 30% between January and May 2026, roughly 6.5 times faster than human traffic over the same window.

Cloudflare CEO Matthew Prince published Radar data on June 3, 2026 showing automated requests had passed humans for the first time: bots generated 57.5% of HTML web traffic, humans 42.5%.

Which AI Crawlers Take the Most and Give Back the Least?

Anthropic and Mistral sit at the extractive end of the crawl-to-refer scale. Google and DuckDuckGo sit near parity. Here is how the June 2026 Cloudflare Radar figures line up:

OperatorPages crawled per referral sentModel
Anthropic (ClaudeBot)~4,580 : 1Answer in place
OpenAI (GPTBot)~848 : 1Answer in place, partial search
Perplexity~186 : 1Cited answer
Google (Googlebot)~5 : 1Index and click
DuckDuckGo (DuckAssistBot)~1.5 : 1Index and click

SEOmator's 2026 GEO Data Report, working from the same Cloudflare source over a different window, ranks Mistral first at 3,389:1, Anthropic second at 2,237:1, and OpenAI at 217:1. The ordering of the extractive tier shifts between reports. The size of the gap against Google does not.

Cloudflare's verified bot share data adds operator scale to the picture. Google runs 28.4% of verified bot traffic, Anthropic 13.2%, Meta 12.2%, and OpenAI 7.2%.

Why Did Meta Become the Largest AI Crawler in Q2 2026?

Meta's two crawlers grew faster than anyone else's and now carry the majority of AI agent traffic on DataDome's network, a position they did not hold in Q1 2026.

Meta-ExternalAgent request volume rose 74% quarter over quarter. Meta-WebIndexer rose 163%. Cloudflare's verified bot data puts Meta at 12.2% of all verified bot traffic, close behind Anthropic's 13.2%.

Meta sends publishers almost no referral traffic in return, so its crawl registers as origin cost with no visible payback. Most B2B marketing teams do not track Meta AI as a discovery surface at all, which leaves the largest crawler on their logs entirely unmeasured.

The crawler leaderboard and the referral leaderboard have now separated. ChatGPT-User request volume fell 6% quarter over quarter while ChatGPT referral traffic grew 17% and held 80% to 88% of all AI referrals every month.

Rank crawlers by referral contribution and citation influence, not by request volume, when you decide what to allow.

How Fast Did AI Crawlers Grow as a Share of Bot Traffic?

AI crawlers went from a rounding error to roughly a quarter of verified bot traffic in about eighteen months.

AI bots averaged 4.2% of HTML requests across 2025. By May 2026, Cloudflare put AI crawlers at 20.3% of verified bot traffic, with AI search bots adding another 6.5%, so AI-related activity reached about 26.7% of the verified total.

HUMAN Security measured agentic traffic growing 7,851% year over year, with automated traffic expanding roughly eight times faster than human traffic.

Operator concentration is high. Google runs 28.4% of verified bot traffic, Anthropic 13.2%, Meta 12.2%, OpenAI 7.2%. Four companies account for more than 60%.

At 750 to 4,000 fetches per day, AI crawling has become a real line item in origin cost and a factor in page performance during burst periods. It also means every blocking decision you make now carries more weight than it did a year ago.

Why Does Google Return Traffic at 5:1 While Anthropic Sits Near 4,580:1?

A crawl returns traffic when the product it feeds sends users to a page. Google indexes your content so a searcher can click a blue link, so the fetch and the visit stay coupled. Anthropic's crawl feeds a model that answers inside Claude, so the fetch and the visit decouple.

The same logic explains Perplexity's mid-range 186:1. Perplexity displays inline sources and a meaningful share of users click them, so some traffic flows back. DuckDuckGo's 1.5:1 reflects a search product that exists to hand off the click.

This reframes what you are optimizing for. Under a 5:1 model, a crawl was an investment that paid in sessions. Under a 4,580:1 model, the crawl pays in influence instead, and you only capture that value if the model names you when a buyer asks who to hire. Nobori's data on AI citation half-life shows how quickly that named position rotates once you win it.

Why Doesn't Any of This Show Up in Google Analytics?

Google Analytics 4 fires on JavaScript execution in a browser session. AI crawlers request the HTML and leave. No script runs, no session opens, no row appears in your reports.

Those requests do land in your server and CDN logs, tagged by user agent: OAI-SearchBot, GPTBot, ClaudeBot, PerplexityBot, Meta-ExternalAgent. Log files are the only place most B2B teams can see the channel at all.

The measurement gap explains a pattern that frustrates marketing leads. Pipeline arrives from buyers who say they found you through ChatGPT, while the analytics dashboard shows flat organic and a shrinking direct channel. The fetches that shaped those buyers' shortlists happened weeks earlier, invisibly, and got attributed to nothing.

Nobori's earlier analysis of the B2B AI visibility gap covers the downstream half of this problem: ranking well in Google while AI engines cite you 3% of the time.

How Much AI Agent Traffic Does a Typical B2B Site Get?

Salespeak tracked more than 640,000 AI agent visits to B2B SaaS websites between January 8 and February 7, 2026. The agents read blog posts, compared vendors and scanned pricing pages. None of it registered as a session in Google Analytics.

Per-site volume in that dataset ran between 750 and 4,000 AI page fetches per day, comparable to what many mid-market B2B SaaS companies see from organic search. ChatGPT accounted for 91% of the AI traffic observed.

Digital Applied's log study across 35-plus sites, including a 12,000-page B2B SaaS documentation property, puts per-crawler daily volume at roughly 4,200 hits per site for GPTBot, 1,800 for ClaudeBot and 980 for PerplexityBot. Combined, GPTBot, ClaudeBot, AppleBot and PerplexityBot generated close to 1.3 billion fetches, about 28% of Googlebot's total volume.

Run the arithmetic on your own domain before you assume you are an exception.

Which Pages Do AI Crawlers Fetch Most on B2B Sites?

They concentrate on comparison, pricing and documentation pages, the same content a buyer reads while building a shortlist.

Digital Applied's crawl-budget audits across B2B properties find AI answers pulling disproportionately from "best X" roundups, "X vs Y" comparisons, alternatives pages, integration documentation, pricing explanations, definitions, and benchmark or FAQ-heavy pages. Salespeak's agent logs show the same shape: blog posts, vendor comparisons, pricing pages.

Crawl allocation spent on thin blog archives returns nothing. A 12,000-page documentation property can starve its own commercial pages when crawlers burn their fetch budget on low-value URLs, which is how sites with strong content end up uncited on the queries that matter.

Rank your URLs by commercial intent, confirm each one gets fetched, then prune or noindex the archive pages competing for the same allocation.

Which OpenAI Crawler Actually Controls Your ChatGPT Citations?

OAI-SearchBot controls them. OpenAI runs three separate user agents with three separate jobs, and confusing them costs B2B teams their ChatGPT visibility:

  • GPTBot collects content for model training. Blocking it opts you out of training and changes nothing about whether ChatGPT cites you.
  • OAI-SearchBot builds the search index ChatGPT reads when it answers with links. Blocking it removes you from ChatGPT search answers.
  • ChatGPT-User fetches live when a person pastes your URL into a chat or a custom GPT calls your page.

OpenAI split GPTBot from its search agents in late 2024. Every robots.txt written during the 2023 blanket-block wave predates that split, which means a rule intended to protect training data now blocks citations too.

DataDome's Q2 data records ChatGPT-User request volume falling 6% quarter-over-quarter while ChatGPT referral traffic grew 17%. Fewer live fetches, more cited answers. The index is doing the work now.

How Many B2B Sites Are Blocking the Crawlers That Cite Them?

CapstonAI's Q1 2026 cohort audit found 41% of B2B sites still blocking at least one major AI bot, most of it leftover configuration from the 2023 and 2024 block-everything period. A separate estimate puts roughly 27% of B2B SaaS and ecommerce sites blocking major LLM crawlers at the CDN layer, often without the marketing team knowing.

Each blocked bot costs an estimated 18% to 34% of your potential citations on that engine. A single stale Disallow line can remove you from an entire platform's answer set while your rankings stay untouched.

Check three layers, not one:

  1. robots.txt for explicit AI user-agent disallows
  2. CDN or WAF bot rules that block by category rather than by name
  3. Rate limiting that returns 429s to legitimate crawlers at their fetch rates

See which engines can reach you and which cannot → nobori.ai

What Happens to Your Site on September 15, 2026?

Cloudflare will block mixed-use AI crawlers by default on any page that hosts ads. The company announced the policy on July 1, 2026, giving AI operators ten weeks to separate their search bots from their training and agent crawlers.

The defaults apply to new Cloudflare customers, new sites added by existing customers, and all existing free-tier customers. Paying customers keep the ability to override from the dashboard and readmit specific crawlers.

Alongside the block, Cloudflare is converting Pay Per Crawl into Pay Per Use, a model that pays publishers when their content shapes an AI answer rather than when a bot fetches the page. Ceramic.ai and You.com signed on as the first two partners. Cloudflare's AI Crawl Control now returns roughly one billion HTTP 402 responses per day.

Two questions for your team this month: does your site sit on a Cloudflare free plan, and do your highest-value pages carry ad units?

Does Web Bot Auth Change Who Gets Into Your Site?

Web Bot Auth replaces spoofable user agents with cryptographic proof. A bot signs each request using HTTP Message Signatures (RFC 9421) with an Ed25519 key, adds a Signature-Agent header, and publishes its keys in a JWKS directory you can verify.

Cloudflare, Amazon, Akamai and OpenAI back the standard, and an IETF working group took it up in 2026. Standards-track specifications went to the IESG in April 2026, with a best-current-practice document on key management due in August 2026 and RFC publication possible in 2027.

Google signs as agent.bot.goog and OpenAI's ChatGPT agent signs as chatgpt.com. Anthropic, Perplexity and Mistral have not documented support.

The practical effect for B2B sites: you can allow verified agents into pricing pages, product specs and comparison content while keeping unverified scrapers out. That turns crawler access from a binary into a policy you control.

What Is an Agent Trust Policy, and Do You Need One?

An agent trust policy is a written rule set defining which automated clients reach which parts of your site. DataDome reports 54% of its customers had adopted one by the end of Q2 2026.

A usable policy specifies five things:

  • Which named agents you allow, path by path
  • What verified agents can reach that unverified ones cannot
  • Rate limits per agent class
  • Which status code unverified traffic receives (402, 403 or 429)
  • Who reviews the allowlist, and how often

Without a policy, your crawler defaults get set by whoever last edited the WAF rules, and marketing discovers the change when citations drop six weeks later. HUMAN Security recorded blocking rates climbing toward 9% by May 2026, much of it unintentional.

Write the policy down, and give the marketing team a review seat before the rules ship.

Which Agentic Browsers Send the Most Real Users?

Browser-based agents carried about 71% of observed agentic activity across the top ten agents in April 2026, according to HUMAN Security's Satori team, led by Perplexity's Comet and OpenAI's Atlas.

The June 2026 rankings: Comet at 47.6% of agentic traffic, Claude at 20.8%, Atlas at 16.5%. In May, Comet held 47% and Atlas 20.3%, while total agentic traffic dipped 4.3% month over month and blocking rates climbed toward 9%.

These agents behave like buyers in the research phase. They click links, fill forms, compare products and navigate multi-step flows. Only 3.16% of agentic activity touches checkout or payment routes.

Industry concentration is extreme. Media took 45.62% of April's agentic traffic, ecommerce 38.20%, travel 14.12%. Those three absorbed 98% of it, which tells B2B teams the wave has not fully arrived in their vertical yet.

How Should B2B Teams Measure Crawl-to-Refer for Their Own Domain?

Pull 30 days of server or CDN logs, count fetches by AI user agent, then divide by referral sessions from each matching platform. You now have a per-operator crawl-to-refer ratio for your own site rather than an industry average.

Four numbers worth tracking monthly:

  1. Fetches by agent. Separate OAI-SearchBot from GPTBot. One predicts citations, the other does not.
  2. Fetch coverage. The share of your indexable URLs that AI crawlers reached at least once. Gaps here become citation gaps.
  3. Response codes by agent. A rising 403 or 429 rate against OAI-SearchBot signals a self-inflicted visibility problem.
  4. Citations per thousand fetches. Your conversion rate from crawl to named mention. This is the number to move.

Fetch counts tell you whether engines can read you. Citation counts tell you whether they choose you. Nobori's study on which page types earn citations covers where to point that crawl budget.

Why Do Crawl-to-Refer Numbers Differ So Much Between Reports?

Three variables move the figure: the measurement window, the denominator, and the level of aggregation. Read every crawl-to-refer number with all three in hand.

Cloudflare Radar's Anthropic ratio ran near 23,951:1 across January to March 2026, 10,300:1 on May 31, and about 4,580:1 in June. The bot did not become four times more generous in a month. Anthropic's referral volume grew off a small base while its crawl rate held, and shorter windows amplify that.

Denominators differ too. Cloudflare reported bots at 57.5% of HTML traffic in June 2026 and 35.2% of all web traffic. Both are correct. HTML requests exclude images, scripts and API calls, which humans generate at far higher rates.

Aggregation is the third trap. Operator-level ratios (Anthropic) and user-agent-level ratios (ClaudeBot) diverge when a company runs several crawlers. Compare like with like or the trend line lies to you.

What Should Your Team Do in the Next 30 Days?

Audit crawler access first. Everything else depends on engines being able to read the pages you want cited.

  1. Week 1. Pull 30 days of logs. Count fetches and response codes for OAI-SearchBot, GPTBot, ClaudeBot, PerplexityBot and Meta-ExternalAgent. Flag any non-200 pattern.
  2. Week 1. Read your robots.txt line by line. Allow OAI-SearchBot and PerplexityBot. Decide on GPTBot deliberately rather than by inheritance.
  3. Week 2. Check your CDN bot rules and your Cloudflare plan tier. Free-tier sites face default blocking on September 15.
  4. Week 2. Calculate your own crawl-to-refer ratio per operator. Record it as a baseline.
  5. Week 3. Map fetch coverage against your commercial pages: pricing, comparisons, alternatives, integration docs. Fix the gaps.
  6. Week 4. Set citations per thousand fetches as the metric your team reports, and review it monthly.

Where Is Crawler Economics Heading?

Two forces are pulling in opposite directions. Crawl volume keeps compounding, with DataDome recording 45% quarterly growth and Fastly measuring AI requests rising 6.5 times faster than human traffic. Access keeps tightening, with Cloudflare defaulting to blocks on September 15 and blocking rates already near 9%.

The likely outcome is a tiered web. Verified agents that identify themselves through Web Bot Auth get broad access. Unverified scrapers get 402s and 403s. Publishers who join Pay Per Use collect on influence rather than clicks.

For B2B teams, the strategic read is narrower and more useful. Referral traffic from AI platforms will stay thin because the ratios are structural, not temporary. Your return on content comes from being the vendor a model names, which makes citation share the metric that matters and crawler access the precondition for earning any of it.

Frequently Asked Questions About the Crawl-to-Refer Ratio

What is a good crawl-to-refer ratio? No universal benchmark exists, because the ratio depends on which platforms crawl you. Compare your own ratio per operator against the June 2026 Cloudflare baselines: Google near 5:1, Perplexity near 186:1, OpenAI near 848:1, Anthropic near 4,580:1. Deviating far above your operator's baseline suggests you are being crawled without being cited.

Should I block AI crawlers to protect my content? Blocking GPTBot opts you out of OpenAI model training without affecting ChatGPT citations. Blocking OAI-SearchBot removes you from ChatGPT search answers. Each blocked bot costs an estimated 18% to 34% of your potential citations on that engine, so decide per user agent rather than per company.

Can I see AI crawler traffic in Google Analytics? No. GA4 requires JavaScript execution in a browser session, and AI crawlers request HTML without running scripts. Server logs, CDN logs and AI visibility platforms are the only places this traffic appears.

Does Cloudflare's September 15 change affect B2B SaaS sites? It affects new Cloudflare customers, new sites added by existing customers, and all free-tier customers, and it applies to pages hosting ads. Most B2B SaaS marketing sites carry no ad units, which narrows the exposure. Verify your plan tier and page inventory rather than assuming.

How often should I audit crawler access? Monthly for response codes and fetch coverage, and immediately after any CDN, WAF or robots.txt change. Blocking regressions ship silently and cost citations for weeks before anyone notices.


Ready to see where you stand in AI search?

Nobori tracks your brand's visibility across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Claude, updated daily. See who's getting cited, where you're missing, and what to fix.

Get daily AI visibility alerts for your brand → nobori.ai

Sources

  • DataDome, AI Traffic Report Q2 2026 (July 16, 2026)
  • Cloudflare Radar, crawl-to-refer and verified bot data (May and June 2026)
  • SEOmator, GEO Data Report 2026: crawl-to-refer ratios
  • HUMAN Security Satori Threat Intelligence, State of Agentic Traffic (April, May, June 2026)
  • Salespeak, AI agent analytics dataset (January 8 to February 7, 2026)
  • Digital Applied, agentic crawler 30-day site log study and AI crawler statistics (2026)
  • Fastly, AI traffic growth analysis (January to May 2026)
  • Cloudflare, AI Crawl Control and Pay Per Use policy announcement (July 1, 2026)
  • IETF Web Bot Auth working group and RFC 9421 HTTP Message Signatures
  • CapstonAI, Q1 2026 B2B cohort robots.txt audit

Recent blogs