When AI Exceeds Human Experts: Measuring Superhuman AI Capabilities | Samira Nedungadi (SecureBio)
Samira Nedungadi (SecureBio) on measuring AI's scientific capability in biology now that models match or exceed human experts on key benchmarks. Nedungadi, who leads engineering at SecureBio, describes how the group benchmarks frontier AI in biology, from knowledge tests like the Virology Capabilities Test to agentic benchmarks that measure what a model can do. The problem she raises: recent models already reach or pass human experts on many of these, so evaluations have to keep pace with super-expert capability. Her approach sources tasks from the leading edge of research and lets reality, not expert consensus, set the correct answer. She introduces two benchmarks: ReproBAIT, where coding agents reproduce published bio-AI tools end to end, and Predictive Bio Bench, developed with NIST, which scores models on predicting outcomes of unpublished experiments. Chapters 0:00 Measuring AI capability in biology (SecureBio) 0:18 AI benchmarks and the Virology Capabilities Test 0:51 Agentic benchmarks: what models can actually do 1:11 Models now match or exceed human experts 1:55 What does beating the experts mean? 2:06 The challenge of measuring super-expert capability 2:43 Two approaches: leading-edge tasks and letting reality decide 3:02 ReproBAIT: reproducing bio-AI tools 4:09 Predictive Bio Bench (with NIST) 4:44 Summary: a diverse panel of estimators More AI safety research: https://far.ai Alignment Workshop playlist: https://youtube.com/playlist?list=PLBY5kyt_LfFg&si=MqopGuY7Koc7dKTc FAR.AI is a research nonprofit working to ensure the safe development of advanced AI. We host the Alignment Workshop series and publish frontier alignment research.