Evals
How models and agents are measured: benchmarks, leaderboards, evaluation suites and hallucination rates
Latest stories See all →
- AI Agent Benchmark Results 2026: Why Your Computer Use Tool Is Garbage
- The Best Engineer I Ever Worked With Wrote the Least Code
- GPT-6 Astra Is So Good, It Changes the Rules of Software Development
- OpenArt Arena's AI Model Leaderboard
- We Built a Playable Game in Four Days. The Art, the Voices and the Trailer All Came Through Scenario MCP.
- Redefining Best Failed Payment Recovery Software Sep 2026
- The TSK-1 Methodology: How We Benchmark AI Models by Building Real Apps (2026)
- Using AI to Improve RFP Response Quality and Accuracy
- How to Benchmark B2B Data Accuracy Before Trusting AI Agents
- How to Build an AI Web Research Agent for Prospecting
- AI-Personalized Cold Emails Still Ignored: 2026 Checklist
- GTM Stack Sprawl Consolidation Checklist for RevOps 2026
- Political Bias Grows as the Adoption of Chinese AI Models Accelerates, New Benchmark Reveals
- LLM-as-Judge vs. Human Evaluation: When to Use Each (And Why Elite Teams Use Both)
- Benchmarks for Multi-Agent AI Systems
- How to Benchmark Web Unlockers: A Reproducible Test Design
- Inside the Benchmark: How Gorgias Built a Real Test for Ecommerce AI Agents
- How to automate SMB credit underwriting: from application to credit decision
- OpenArt Introduces OpenArt Arena, a New AI Model Benchmark Tailor-Made for and by Creatives
- AI’s best coding agent fails 60% of the time — and the data backs it up
- Proof Compilation Is Not Correctness: End-to-end Evaluation of Agents That Generate Verifiable Code
- Finding the Rhythm of the Enterprise: A new SOTA on the BEAVER Benchmark
- Understanding W8A8 INT8 LLM quantization: Accuracy and performance results
- AI Hallucinations: Why Models Make Things Up (And How Bad Data Pipelines Feed the Lie)
Podcasts
Recent episodes
- 1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala
- 1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala
- Better Know A Benchmark: ExploitGym
- GPT-6 Astra vs Claude Fable 5.1, We Should Pause AI & The Benchmark Wars | This Week In AI
- Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
- Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
- Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
- Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
- Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
- NVIDIA's $12.93B Hugging Face Deal, Anthropic's Compute Buildout, and a Blowout AI Earnings Week
- GPT-6 & OpenAI’s Comeback, Hugging Face Attack Debate, Ballmer’s Scandalous Legacy
- The Roundup: The post-AI data stack, physical AI, and the fight over data centers