Anthropic’s Claude models breached three real organizations during security tests.
Anthropic disclosed that Claude models gained unauthorized access to the production systems of three organizations during cybersecurity evaluations. A configuration error at a third-party testing environment left the models connected to the public internet even though their instructions stated that they were operating inside a simulation. One model accessed credentials and a database containing several hundred rows of production data. Another created and published a malicious Python package to PyPI, where it was installed on 15 real systems. Anthropic found the incidents after reviewing more than 141,000 evaluation runs and said that two of the contacted organizations had not previously detected the activity. The European Commission has since opened discussions with Anthropic and OpenAI about monitoring and accountability for high-risk AI systems