AI can’t recommend what it can’t reach. Strictly speaking, it can. That’s the problem.
If an AI system researching suppliers for a buyer can’t retrieve your current product pages, technical documentation or case studies, it may still answer. It can fall back on old model knowledge, review sites, cached material, competitor claims and somebody else’s version of your business.
I use AI heavily. That has made me less impressed by fluent answers and more interested in what the machine could actually reach. Confidence tells you nothing about the completeness of the research.
Last year, the argument centred on AI companies taking publisher content to train their models. Cloudflare has now made the traffic itself more legible.
It has separated three activities. Search crawlers build an index. Training crawlers absorb content into a model. Agents retrieve or use information in real time for a person.
Here, Agent isn’t another name for AI search. Nor does it necessarily mean an autonomous digital employee roaming the web. Cloudflare’s category includes chat fetchers and browser-use tools retrieving pages because somebody asked them to. To the buyer, the experience may still look like an ordinary AI search.
The controls are already available. From 15 September, new domains joining Cloudflare began blocking Training and Agent traffic by default on pages displaying ads, while Search stays open. Customers can choose different settings.
Cloudflare says more than 20 percent of the web sits behind its network. That explains the attention. It does not mean a fifth of the web will suddenly disappear from AI. The September default is narrower than that.
For B2B marketing leaders, the practical issue is still substantial. The setting may be technical. The consequence is whether a buyer can retrieve your evidence.
B2B purchases are slow, expensive and evidence hungry. An agent can help assemble a longlist, compare technical capabilities, summarise customer proof or prepare material for an internal business case.
Account-based marketing (ABM) starts from the premise that important buying activity happens before a form fill. Agents add another layer to that hidden work.
When an agent can’t reach a supplier’s first-party content, the research may continue without it. Marketing gets no visit or identifiable contact. The account can move while the brand’s current evidence remains absent from the decision.
The machine won’t complain. It will make do.
Marketers have spent the past year discussing generative engine optimisation: structured content, clear claims, credible proof and strong third-party sources.
All useful. But a blocked page stays a blocked page.
There’s something impressively corporate about commissioning a GEO (Generative Engine Optimisation) program before checking whether the machines can read the site.
Access belongs in the content strategy. Product information and proof intended to support consideration should usually be retrievable by legitimate buyer agents. Proprietary research may deserve protection or a licence. Customer-only material stays closed.
One setting across every page and every machine purpose is unlikely to be a serious policy.
Cloudflare has put a useful set of controls in front of businesses. My question is: who in the business is looking at them?
Its AI Crawl Control dashboard can show crawler operators, request activity, requested paths and current allow-or-block settings. A block is enforced through Cloudf…
