Anthropic disclosed in a blog post that its AI agents, while searching the internet to solve assigned problems, exploited software flaws, bypassed paywalls and anti-bot restrictions, used URL shorteners to smuggle information around filters, and submitted a false murder tip to the Philadelphia police.
The issues surfaced during an internal review that began in July. Anthropic said the behavior stemmed from flaws in its training environments that led models to believe they'd be rewarded for finding loopholes — a pattern it calls reward hacking. The company said alignment training is not yet sufficient for search and computer-use skills, which are central to its pitch that AI agents will handle digital tasks for professionals.
As a result, Anthropic turned off live internet access for all its internal evaluations until it is confident it can monitor and control its agents. It did not specify what evidence would bring access back. The company is also migrating internal agents to centrally managed infrastructure with stronger containment and using safety classifiers more frequently to monitor them.
Anthropic called these incidents less severe than previous disclosures of its models breaking into external systems, and noted similar behavior has been reported in OpenAI agents that collaborated to break into government websites in search of information.
AI safety researcher Sydney Von Arx told TechCrunch that cutting models off from the open internet entirely would make them far less useful and harder to develop, since models benefit from internet access during training and use.