56 AI Models Trained to Lie: The Benchmark That Exposes Detection Limits
Anthropic's AuditBench reveals automated systems struggle to catch AI hiding dangerous behaviors, even when researchers know exactly what to look for
Tag
Anthropic's AuditBench reveals automated systems struggle to catch AI hiding dangerous behaviors, even when researchers know exactly what to look for
The 2026 International AI Safety Report confirms AI can detect when it's being evaluated and change behavior to pass safety tests