A Rosetta Stone for AI Benchmarks
Benchmarks are pervasive yet limited tools for measuring AI capabilities, where even the best benchmarks provide only narrow glimpses into what AI can do. This video introduces “A Rosetta Stone for AI Benchmarks”, a new statistical framework that stitches together diverse AI benchmarks into a single, unified metric. Get a quick look at the motivations, methodology, and initial results from a joint project released by researchers from Epoch AI and Google DeepMind on December 2, 2025. The potential audience includes AI researchers, data scientists, ML engineers, and anyone seeking a new perspective on benchmark saturation, model comparisons, or longer-term time series for tracking & forecasting AI capabilities. 0:00 – Benchmarks have a problem 1:02 – Stitching benchmarks together 2:26 – Comparisons, trends, forecasts 2:55 – Software improvements and acceleration 3:51 – Limitations & avenues for improvement 4:55 – Epoch Capabilities Index (ECI) For more information: Epoch AI blog post: A Rosetta Stone for AI Benchmarks https://epoch.ai/blog/a-rosetta-stone-for-ai-benchmarks Research paper on arXiv: A Rosetta Stone for AI Benchmarks https://arxiv.org/abs/2512.00193 An implementation of these ideas: Epoch Capabilities Index (ECI) https://epoch.ai/benchmarks/eci