AI Crawl to Referral: The Energy B2B Measurement Gap
AI crawlers fetch energy B2B pages at industrial scale and send back almost nothing. Cloudflare Radar data for June 2026, as reported in trade analysis, puts Anthropic's crawl to referral ratio near 4,580 to 1 and OpenAI's near 848 to 1 against Googlebot's 5 to 1. Referral counting is therefore the wrong instrument. The chain to measure is fetch, then citation, then referral, and most energy marketing teams are only looking at the last link.
How should energy B2B teams measure AI search visibility? Not by referral traffic alone, because AI assistants fetch far more than they send. Cloudflare Radar data for June 2026, as compiled in trade analysis, puts the crawl to referral ratio at roughly 4,580 to 1 for Anthropic's crawler, 848 to 1 for OpenAI's, 186 to 1 for Perplexity, 5 to 1 for Googlebot and 1.5 to 1 for DuckDuckGo. A page can be read thousands of times and generate a single visit, so a referral based dashboard will report near zero and conclude wrongly that nothing is happening. The measurable chain has three links: fetches by agent from server logs, citations in assistant answers for your real buying questions, and referrals in analytics. Instrument all three, and report citations per thousand fetches as the conversion metric between the first two.
- The ratios are order of magnitude indicators, not metrics. Cloudflare Radar for June 2026 as reported in trade analysis gives roughly 4,580 to 1 for Anthropic, 848 to 1 for OpenAI, 186 to 1 for Perplexity, 5 to 1 for Googlebot and 1.5 to 1 for DuckDuckGo. A separate dataset, SEOmator's 2026 GEO Data Report, ranks Mistral at 3,389 to 1 and Anthropic at 2,237 to 1 over a different window. The disagreement is the lesson.
- Fetch volume against an energy B2B site is not trivial. Salespeak, tracking B2B SaaS properties between 8 January and 7 February 2026, reported per-site volumes between 750 and 4,000 AI page fetches a day, with ChatGPT accounting for 91 per cent of observed AI traffic. Digital Applied's analysis across more than 35 sites put combined crawler activity at about 1.3 billion fetches, roughly 28 per cent of Googlebot's volume.
- A large share of B2B sites are still blocking the crawlers they want to be cited by. CapstonAI's Q1 2026 audit reported 41 per cent of B2B sites blocking at least one major AI bot. In energy specifically, where legal and security teams often own robots policy, the block is frequently inherited rather than decided.
- Separate the agents before you count them. OpenAI operates distinct crawlers for training and for live answer retrieval, and only the retrieval agent is in the path that produces a citation. Aggregating them into one AI bot row destroys the only signal that predicts visibility.
- Citations per thousand fetches is the metric worth owning. It converts a volume you cannot control into a conversion rate you can improve, and it is the only number in this chain that moves when you change the page rather than when the vendor changes its crawl budget.
- The referral number is still worth collecting, with low expectations. Treat it as confirmation that the chain completed, not as the measure of AI visibility, and expect the long energy sales cycle to break attribution regardless, which is a problem we have covered separately.
Because the economics of an assistant answer remove the click
An assistant that answers a buying question well has no reason to send a visit. It read your page, extracted the specification, the figure or the definition it needed, and returned it inside its own interface. The user got the answer. The citation may be a footnote the user never opens. That is not a failure of your content, it is the product working as intended, and it means referral traffic measures the residue of the interaction rather than the interaction.
The scale of the asymmetry is what makes referral counting actively misleading rather than merely incomplete. Cloudflare Radar data for June 2026, as compiled in trade analysis, puts the crawl to referral ratio at approximately 4,580 to 1 for Anthropic's crawler and 848 to 1 for OpenAI's, against 5 to 1 for Googlebot and 1.5 to 1 for DuckDuckGo. Perplexity sits between them at about 186 to 1. A site being read heavily by an assistant and a site being ignored by it look almost identical in a referral report.
Those figures should be handled as orders of magnitude rather than as measurements. SEOmator's 2026 GEO Data Report, over a different window, ranks Mistral at 3,389 to 1 and Anthropic at 2,237 to 1, roughly half the Cloudflare derived figure for the same operator. Both cannot be precise. The honest use of either is to establish that the ratio is in the hundreds to thousands, which is enough to disqualify referral traffic as the primary instrument and not enough to benchmark against.
This is the same error we documented when search impressions were being read as buyer demand. In both cases a number that is easy to retrieve is standing in for a number that is hard to retrieve, and the substitution is never declared.
Enough that it is a server log question, not a curiosity
The volumes reported in B2B contexts are substantial. Salespeak, tracking B2B SaaS properties between 8 January and 7 February 2026, found per-site volumes running between 750 and 4,000 AI page fetches per day, with ChatGPT accounting for 91 per cent of the AI traffic observed in that sample. Digital Applied, analysing more than 35 sites, put combined AI crawler activity at approximately 1.3 billion fetches, about 28 per cent of Googlebot's total volume over the same scope.
Energy B2B has a specific profile inside those numbers. The content that assistants pull hardest is exactly the content energy sellers produce most of: specification detail, standards comparisons, procurement definitions, regulatory timelines and figure-bearing market explainers. These are high-extraction documents. A buyer asking an assistant which standard applies to a wellhead specification, or when a compliance regime starts, is asking a question your technical content answers completely, which maximises the probability of extraction and minimises the probability of a visit.
The counterpart finding is that a large share of sites are shut to this traffic without having decided to be. CapstonAI's Q1 2026 audit reported 41 per cent of B2B sites still blocking at least one major AI bot. In energy the robots policy is often set by a security or legal function with a default-deny posture, written before the retrieval crawlers existed, and never revisited by the marketing team that now depends on them. Our note on AI crawler access and llms.txt covers the mechanics of auditing and correcting that.
The practical first step is unglamorous. Pull the server or CDN logs, filter by user agent, and establish whether you are being fetched at all. A surprising number of teams have been optimising for AI visibility for a year without ever confirming that the crawler can reach the page.
Three links in the chain, and a conversion rate between the first two
The chain is fetch, then citation, then referral. Each link has a different owner, a different data source and a different failure mode, and collapsing them is what produces the reporting that satisfies nobody.
Fetch is a server log question. Count fetches by agent monthly, and separate the agents properly, because OpenAI operates distinct crawlers for training corpus collection and for live answer retrieval and only the retrieval agent sits in the citation path. Measure fetch coverage as the share of your indexable URLs that were fetched at all, which exposes the pages the assistant has never seen. Measure response codes by agent, because a crawler receiving a high rate of 403 or 429 responses is being blocked by rate limiting or bot management that nobody intended to apply to it.
Citation is a sampling question and the one most teams skip because it cannot be automated cheaply. Build a fixed panel of the real questions your buyers ask, phrased as a buyer would phrase them rather than as a keyword, run that panel across the assistants your market uses on a fixed cadence, and record whether you were cited and in what position. The panel must be stable across periods or the series is meaningless. Thirty to fifty questions is enough to see movement in energy B2B, where the question space is narrower and more technical than in consumer categories.
Referral is an analytics question with a known defect. Assistant referrers are inconsistent, often stripped, and in-app browsers frequently present as direct traffic, so the referral number is a floor rather than a measurement. Collect it, label it as a floor, and do not build a target on it.
The metric to own across the first two links is citations per thousand fetches. It is a conversion rate. Fetch volume is set by the vendor's crawl budget and is largely outside your control. Citation rate responds to how answerable, how specific and how well structured the page is, which is entirely inside your control. A rising citation rate on flat fetch volume is the signature of content work that is landing.
Fetch · Server or CDN logs · Fetches by agent, fetch coverage of indexable URLs, response codes by agent · Partly, via access and performance
Citation · Fixed question panel, run manually or by tool · Citation rate and position across a stable panel of buyer questions · Yes, this is the work
Referral · Web analytics · Sessions from assistant sources, labelled as a floor not a measure · No
Conversion · Derived · Citations per thousand fetches · Yes, and this is the metric to target
Answerability, specificity and a figure that can be lifted cleanly
A retrieval system is selecting a passage it can quote with confidence. That selection rewards properties that are mostly unrelated to classic ranking. A direct answer in the first two to four sentences, placed before the context and the narrative, is the single highest leverage change, because it gives the model a complete and quotable unit. A question used as a literal heading matches the user's phrasing more closely than a keyword headline does. A figure with a named source and a date attached is far more likely to be lifted than the same figure stated bare, because the model can carry the attribution with it.
The inverse also holds and is worth stating because it contradicts a lot of existing practice. Content engineered to withhold the answer in order to force a click performs badly here. It gives the retrieval system nothing to extract, so it is neither cited nor visited. The gated asset is the extreme case: a page that promises the answer behind a form is invisible to the entire chain.
Energy B2B has a structural advantage it under-uses. The sector's buying questions have correct answers. Which standard governs this component, when does this regulation bind, what is the typical lead time, how is this score calculated. Correct, specific, sourced answers to narrow technical questions are precisely what retrieval systems prefer, and most energy vendors still publish capability narrative instead. The opportunity is to write the reference document for the questions your buyers actually ask, which is the same discipline that produced our standing work on answer engine optimisation.
One caution. None of this substitutes for commercial qualification. A citation is reach, not pipeline, and the energy sales cycle is long enough that the connection between the two will not appear in a quarterly report. Treat citation rate as a leading indicator and keep the revenue attribution conversation separate, on the long cycle terms we have set out elsewhere.
One conversion rate, one coverage number, and an honest floor
The reporting failure in this area is usually over-claiming. A dashboard that presents AI referral sessions as the measure of AI visibility will show a number near zero and invite the conclusion that the channel is not worth funding, which the fetch data directly contradicts. A dashboard that presents fetch volume as success over-claims in the other direction, because being crawled four thousand times a day is not an achievement, it is a precondition.
A defensible monthly report has three lines. Citations per thousand fetches, as the conversion metric and the thing the team is accountable for. Fetch coverage, as the share of indexable pages the assistants have actually seen, which exposes technical and access problems early. And referral sessions, labelled explicitly as a floor because referrer data is stripped and in-app browsers misattribute. Add the citation panel result as a trend rather than a snapshot, since position movement across a stable question set is the clearest evidence of progress.
The discipline is the same one that applies to every measurement claim in this business. Name the source, name the window, state the method, and mark an estimate as an estimate. The figures in this article are a case in point: two reputable datasets disagree by a factor of two on the same operator's crawl to referral ratio, and a report that quoted either one as precise would be wrong in a way its author could not detect.