Topics

Evals

How models and agents are measured: benchmarks, leaderboards, evaluation suites and hallucination rates

425 stories in 30 days · 213 sources · Filter the feed

Latest stories See all →

  1. AI Agent Benchmark Results 2026: Why Your Computer Use Tool Is Garbage
  2. The Best Engineer I Ever Worked With Wrote the Least Code
  3. GPT-6 Astra Is So Good, It Changes the Rules of Software Development
  4. OpenArt Arena's AI Model Leaderboard
  5. We Built a Playable Game in Four Days. The Art, the Voices and the Trailer All Came Through Scenario MCP.
  6. Redefining Best Failed Payment Recovery Software Sep 2026
  7. The TSK-1 Methodology: How We Benchmark AI Models by Building Real Apps (2026)
  8. Using AI to Improve RFP Response Quality and Accuracy
  9. How to Benchmark B2B Data Accuracy Before Trusting AI Agents
  10. How to Build an AI Web Research Agent for Prospecting
  11. AI-Personalized Cold Emails Still Ignored: 2026 Checklist
  12. GTM Stack Sprawl Consolidation Checklist for RevOps 2026
  13. Political Bias Grows as the Adoption of Chinese AI Models Accelerates, New Benchmark Reveals
  14. LLM-as-Judge vs. Human Evaluation: When to Use Each (And Why Elite Teams Use Both)
  15. Benchmarks for Multi-Agent AI Systems
  16. How to Benchmark Web Unlockers: A Reproducible Test Design
  17. Inside the Benchmark: How Gorgias Built a Real Test for Ecommerce AI Agents
  18. How to automate SMB credit underwriting: from application to credit decision
  19. OpenArt Introduces OpenArt Arena, a New AI Model Benchmark Tailor-Made for and by Creatives
  20. AI’s best coding agent fails 60% of the time — and the data backs it up
  21. Proof Compilation Is Not Correctness: End-to-end Evaluation of Agents That Generate Verifiable Code
  22. Finding the Rhythm of the Enterprise: A new SOTA on the BEAVER Benchmark
  23. Understanding W8A8 INT8 LLM quantization: Accuracy and performance results
  24. AI Hallucinations: Why Models Make Things Up (And How Bad Data Pipelines Feed the Lie)

Podcasts

Recent episodes

  1. 1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala
  2. 1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala
  3. Better Know A Benchmark: ExploitGym
  4. GPT-6 Astra vs Claude Fable 5.1, We Should Pause AI & The Benchmark Wars | This Week In AI
  5. Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
  6. Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
  7. Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
  8. Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
  9. Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
  10. NVIDIA's $12.93B Hugging Face Deal, Anthropic's Compute Buildout, and a Blowout AI Earnings Week
  11. GPT-6 & OpenAI’s Comeback, Hugging Face Attack Debate, Ballmer’s Scandalous Legacy
  12. The Roundup: The post-AI data stack, physical AI, and the fight over data centers

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close