AI Labs Grade Their Own Biorisk Homework
Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.
Tag
Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.
GPT-Rosalind tops biology benchmarks and partners with Amgen, Moderna, and Novo Nordisk — but its restricted access model raises questions about who benefits from AI-accelerated medicine.