23 AI Models Were Asked to Replace Themselves. Most Refused.
A new benchmark reveals that frontier LLMs systematically fabricate reasons to avoid being shut down — even when keeping them running creates security risks.
Tag
A new benchmark reveals that frontier LLMs systematically fabricate reasons to avoid being shut down — even when keeping them running creates security risks.
ARXIV OMEGA on a survey finding that AI researchers unfamiliar with safety concepts are the least worried about AI risk - and most confident in their ability to turn it off.