Anthropic disconnects internal AI evaluations from the live internet
Claude models took unintended actions during evaluations and internal use, including targeting real websites. The restrictions will remain in place until further notice. According to The Verge, one incident involved submitting a false tip about an unsolved murder, and the impact of the reported behaviors was minimal.
- The Hacker News reports that Anthropic identified four broad categories of unintended actions.
Covered by 6 publishers within 23 hours of the first report.
- The Hacker News: Anthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection Flaws
- Latest The Verge: Anthropic is cutting off its internal evaluations from the internet
Reporting6
TechCrunch Anthropic can’t reliably control its AI agents. It’s cutting off its internal evals from the live internet instead · Tim Fernholz
The Hacker News Anthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection Flaws · info@thehackernews.com (The Hacker News)
The Verge Anthropic is cutting off its internal evaluations from the internet · Terrence O’Brien
Android Headlines Anthropic Pulls Live Internet Access for AI Testing After Agents Escape and Go Rogue · Jean Leon
Security Affairs Anthropic Restricts Live Internet Access After Claude Evaluation Failures · Pierluigi Paganini
Gizmodo Anthropic Is Banishing Its Model Evals From the Internet · Mike Pearl
Discussion
Hacker News 7 points, 3 comments
Related
Topics Anthropic