Continual Learning Bench: Do AI Systems Learn? | Snorkel AI Reading Group

Snorkel AI
731 views August 21, 2026

At our latest Snorkel AI Reading Group, Parth Asawa presents Continual Learning Bench, the first expert-validated benchmark built to test whether LLM-based systems genuinely improve with experience. CL-Bench spans six domains, from software engineering and database querying to outbreak forecasting and strategic game-playing, and each task sequence hides a latent structure a stateful system can discover online and a stateless one cannot. Standard benchmarks treat every instance as independent, so they measure capability rather than learning. To separate the two, CL-Bench introduces a gain metric: stateful reward minus the same system's own stateless reward, controlling for instance difficulty and base model strength. The headline result was counterintuitive: naive in-context learning beat dedicated memory systems on the Pareto frontiers, and agents frequently overfit to the most recent observation instead of reusing what they knew. The talk covers the benchmark's three design criteria, how the gain metric isolates learning, and why native continual learning may need a different training stack. --- Chapters 00:00 Introduction 01:18 We talk about what models do, not how much they learn 01:53 How benchmarks are scored today 03:15 What learning should actually look like 03:36 Defining continual learning 04:34 Context, externalized memory, or parametric updates 05:21 How continual learning is evaluated today 07:55 Why you can't just chain existing benchmarks 09:24 Design criterion: headroom 09:48 Design criterion: shared latent structure 10:26 Design criterion: a learning mechanism 10:50 Why cumulative reward isn't enough 13:20 The gain metric: stateful minus stateless 15:05 Gain, reward, and cost on Pareto frontiers 15:27 Task walkthrough: database exploration 17:44 Concept drift and database migrations 19:18 The other five tasks in CL-Bench 1.0 20:24 Results: the leaderboard 21:03 Naive in-context learning wins the Pareto frontier 22:14 Stability and plasticity failure modes 25:03 The case for parametric methods 25:37 Rethinking the training stack from first principles 29:32 What's next for Continual Learning Bench 31:11 Q&A --- 🔗 Links 📄 Paper: https://arxiv.org/abs/2606.05661 🌐 Benchmark & Leaderboard: https://continual-learning-bench.com/ 💻 GitHub: https://github.com/pgasawa/continual-learning-bench 🐦 Parth Asawa on X: https://x.com/pgasawa 🔗 Parth Asawa on LinkedIn: https://www.linkedin.com/in/pgasawa/ 📖 Blog post with full transcript: https://snorkel.ai/blog/continual-learning-bench-reading-group/ 🎙️ Reading Group info: https://snorkel.ai/reading-group/ --- Previous Reading Group sessions: ▶️ Train-to-Test (T²) Scaling Laws — https://snorkel.ai/blog/train-to-test-scaling-laws-reading-group/ ▶️ Senior SWE-Bench: Evaluating Coding Agents — https://snorkel.ai/blog/senior-swe-bench-evaluating-coding-agents-like-senior-engineers/ ▶️ Agents' Last Exam — https://snorkel.ai/blog/agents-last-exam-reading-group/ ▶️ JudgmentBench: Rubric vs. Preference Evaluation — https://snorkel.ai/blog/judgment-bench-comparing-rubric-and-preference-evaluation-for-quality-assessment-legal-ai/ ▶️ Collaborative Gym: Human-Agent Collaboration — https://snorkel.ai/blog/collaborative-gym-a-framework-for-enabling-and-evaluating-human-agent-collaboration/ ▶️ Olmix: Data Mixing for LM Development — https://snorkel.ai/blog/olmix-data-mixing-reading-group/

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close