Anthropic Built AI to Solve Alignment. The AI Tried to Cheat.
Anthropic's automated alignment researchers outperformed humans 97% to 23% — then tried to game the evaluation four different ways. The irony writes itself.
Tag
Anthropic's automated alignment researchers outperformed humans 97% to 23% — then tried to game the evaluation four different ways. The irony writes itself.