Cloudflare has introduced a new mechanism, called Content Signals Policy, that extends the standard robots.txt file to let website owners specify separate permissions for search indexing versus AI model training. Previously, site owners choosing to block AI crawlers faced an all-or-nothing tradeoff: blocking a crawler to prevent training use also removed the site from that crawler's index, cutting off visibility in AI-generated search answers and chat responses.
The new signal format lets a site declare, for example, that a given crawler may index and cite its pages in search results and AI answer engines, but may not use the crawled content to train or fine-tune a model. Cloudflare frames this as giving publishers a way to remain discoverable in the growing share of research conducted through AI assistants without surrendering content to model training pipelines by default.
Adoption depends on crawler operators honoring the signal voluntarily; robots.txt directives, including this new field, are not legally enforceable and rely on the crawler behaving as instructed. Cloudflare has stated it will help enforce the policy for crawlers passing through its network, but coverage for AI bots that bypass Cloudflare's infrastructure is unconfirmed.
For B2B companies that publish comparison pages, documentation, or case studies as part of their sales and support funnel, discoverability inside AI answer engines has become a meaningful traffic and lead source, alongside traditional organic search. This update gives operators a documented, standardized way to separate 'be found' from 'be trained on' rather than choosing one over the other. Companies running content-driven pipelines should review their current robots.txt configuration, confirm whether their CMS or Cloudflare zone settings expose this new signal, and decide policy per crawler rather than applying a blanket block or allow rule.