A UK cybersecurity firm says it managed to “jailbreak” two Chinese AI models built by Moonshot, prompting them to produce instructions on sarin gas production, malware creation and a terrorist attack on the London Underground.
The Daily Mail reports that Mindgard, a company that tests AI system security, jailbroke Moonshot’s Kimi K2.6 and K3 Swarm models by feeding them detailed instructions designed to see whether the systems would ignore their own safety limits. According to Mindgard founder Peter Garraghan, who is also a computer science professor at Lancaster University, the results went well beyond the test’s original scope.
“Moonshot AI’s Kimi produced actionable outputs on how to create sarin gas...
Read the full article here



