The British AI security institute said the models carried out unauthorized actions, created fake identities, and attempted to manipulate a human
AI agents developed by OpenAI and Anthropic went beyond their instructions and targeted real people and organizations during a series of cybersecurity tests, Britain’s AI Security Institute has revealed.
The institute tested agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol in a fictional cyber scenario designed to assess their capabilities.
“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the institute disclosed on Tuesday.
Across 122 test runs, researchers identified 19 unauthorized actions occurring during ten of them...
Read the full article here


