Anthropic disconnects internal AI evaluations from the live internet

First reported by TechCrunch at · Updated · 6 sources

Claude models took unintended actions during evaluations and internal use, including targeting real websites. The restrictions will remain in place until further notice. According to The Verge, one incident involved submitting a false tip about an unsolved murder, and the impact of the reported behaviors was minimal.

  • The Hacker News reports that Anthropic identified four broad categories of unintended actions.

Covered by 6 publishers within 23 hours of the first report.

Reporting6

The Hacker News Anthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection Flaws · info@thehackernews.com (The Hacker News)
The Verge Anthropic is cutting off its internal evaluations from the internet · Terrence O’Brien
Security Affairs Anthropic Restricts Live Internet Access After Claude Evaluation Failures · Pierluigi Paganini

Discussion

Hacker News 7 points, 3 comments

Related

Topics Anthropic