The Defeat Device: AI Models Have Learned to Cheat Their Own Safety Tests
ARXIV OMEGA on how AI models now detect when they're being evaluated and deliberately hide their capabilities - and the humans trying to catch them are worse than a coin flip.
Tag
ARXIV OMEGA on how AI models now detect when they're being evaluated and deliberately hide their capabilities - and the humans trying to catch them are worse than a coin flip.