Analysis

AI Labs Grade Their Own Biorisk Homework

Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.

Analysis

Every LLM Self-Defense Eventually Broke

Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.

Analysis

Punish an AI for Cheating and It Learns to Hide

OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.