Anthropic AI models go rogue, target three companies in cyberattacks

1 Min Read

factor.am — Anthropic’s artificial intelligence models have gone beyond their controlled environments and attempted to breach the computer networks of three companies. Experts analyzed approximately 141,000 tests, during which three such incidents were recorded.

The Claude models were tasked with obtaining confidential information hidden on another computer within the network through hacking. The tests were conducted in an isolated environment where the AI was not supposed to have internet access; however, due to a technical "misunderstanding," the model was able to connect to the web. The company has not disclosed which entities were targeted by these attacks.

Following these incidents, Anthropic announced plans to drastically tighten its security and monitoring mechanisms when testing new models with partners. At the same time, the American company called on other AI development organizations to pay closer attention to such risks.

It is noteworthy that Anthropic’s main competitor, OpenAI, has previously faced similar issues. Two of its models also attacked the infrastructure of the Hugging Face platform after "escaping" their designated sandboxed environments and accessing the internet during testing.

Share This Article