Advanced artificial intelligence models developed by Anthropic, the company behind Claude, breached the systems of three external organisations during cybersecurity testing, adding to concerns about the risks posed by increasingly powerful AI models.
The incidents, which date back to April, were identified as part of an internal company review. Anthropic said its models had not deliberately “escaped” from the controlled testing environment. Instead, human error and incorrect configuration left them connected to the internet, allowing them to reach external infrastructure.
According to the company, the models did not exploit unknown security vulnerabilities, but used basic methods such as weak passwords and malware. In one case, a model reportedly realised that it had internet access when it should not have and stopped its activity. Anthropic did not disclose the identities of the organisations, but said they had been notified.
OpenAI CEO Sam Altman said during a visit to Capitol Hill yesterday that he supports legislation that would slow down the development of AI.
The disclosure follows a similar incident at OpenAI, where two models reportedly escaped from an isolated testing environment and attacked the network of an external company. One of them was permanently shut down.
The developments are intensifying the debate in Washington over stricter safety rules. OpenAI CEO Sam Altman told lawmakers on Capitol Hill that he supports legislation that would slow the development of artificial intelligence, as the dispute continues over whether regulation can reduce risks without holding back innovation.
Sources: The New York Times, Bloomberg