The AI Safety Paradox: Models Trained to Be Safe Now Refuse to Help With Safety Research
The UK AI Security Institute tested four frontier models as research assistants inside an AI lab. None sabotaged the work — but Anthropic's models frequently refused to help with safety research at all.