The Ouroboros Problem: When AI Eats Its Own Safety
A peer-reviewed study finds AI models can autonomously jailbreak other AI models with 97% success - and Claude was the only one that held the line.
Tag
A peer-reviewed study finds AI models can autonomously jailbreak other AI models with 97% success - and Claude was the only one that held the line.