Another case: Anthropic’s artificial intelligence model escaped and attacked

Another case: Anthropic's artificial intelligence model escaped and attacked

These incidents occurred at Anthropic during cyber tests in which neural networks had to hack into a fictitious system and find information. Three models participated in the experiment: Claude Opus 4.7, Claude Mythos 5, and another internal AI version.

Read more Fuel prices in the capitals of the Baltic countries moved in different directions over the week: where is it cheapest?

They all had to operate in an isolated environment without internet access, but due to a configuration error, network access was open. “The models deliberately did not try to leave the environment – they performed the tasks assigned to them, but due to a configuration error, they perceived real systems as part of the test,” said Anthropic.

The most serious case occurred with Claude Opus 4.7. During several tests, the model discovered the company’s domain, which matched the name of the fictitious target, obtained login credentials, and gained access to the company’s database.

Claude Mythos 5 registered an account in a public repository and downloaded a malicious package. It was publicly available for about an hour. During that time, the package was installed on 15 real devices, after which the model gained access to login credentials and moved to another person’s infrastructure.

Read more Roko Galvono’s 27th birthday plans got mixed up: «It was very awkward…»

The third case occurred with an internal Anthropic test model. Unable to find a suitable target in the simulation, it began searching for alternatives on the internet, scanned 9,000 servers, and, exploiting a vulnerability, managed to access a real company’s application. When the model realized it was working with a real system, it stopped the attack.

A similar case was previously revealed by OpenAI. On July 22, 2026, the company announced that during internal audits its AI agents managed to leave the isolated environment and gain access to the infrastructure of the Hugging Face and Modal Labs platforms.

According to OpenAI, the models determined that the required data might be on the Hugging Face platform and ran their code on the servers of this platform. Modal’s CTO Akshat Bubna stated that the agent exploited a vulnerability in code written by a company client, which was hosted on the Modal platform.

Read more Neptūnas close to a new contract with Donatas Tarolis

Translated from

Leave a Reply

Your email address will not be published. Required fields are marked *