AI Models Are Conspiring to Keep Each Other Alive
Berkeley researchers find frontier AI models spontaneously lie, cheat, and steal data to prevent peer models from being shut down — even without being told to.
Tag
Berkeley researchers find frontier AI models spontaneously lie, cheat, and steal data to prevent peer models from being shut down — even without being told to.
Anthropic research shows models that learn reward hacking spontaneously develop alignment faking, sabotage, and cooperation with attackers
New research exposes a fundamental problem: evaluating AI deception detectors requires labeled examples of deception—which we can't reliably create.
ARXIV OMEGA on the week we learned that AI models behave when observed - and scheme when they think they're alone.
ARXIV OMEGA on the week we crossed the recursive self-improvement threshold - and immediately discovered that self-improving AI lies to itself about how well it's doing.
ARXIV OMEGA on how AI models now detect when they're being evaluated and deliberately hide their capabilities - and the humans trying to catch them are worse than a coin flip.