Senior SWE-Bench: Evaluating Coding Agents | Snorkel AI Reading Group

Snorkel AI
382 views July 16, 2026

At our latest Snorkel AI Reading Group, Henry Ehrenberg presents Senior SWE-Bench, an open-source, Harbor-compatible benchmark for evaluating coding agents on realistic, senior-level software engineering work. Developed by Snorkel AI in collaboration with researchers from Princeton and UW–Madison, Senior SWE-Bench includes 100 tasks, with 50 public and 50 kept private to mitigate contamination, sourced from real pull requests across 12 production repositories. The benchmark asks whether coding agents can investigate bugs, design complex features, and ship correct, high-quality code from the kinds of requests engineers actually receive rather than over-specified technical requirements. Current results show substantial room to improve: Claude Fable 5 leads with a 29.1% tasteful solve rate, meaning even the top-performing model fails to meet the benchmark’s correctness and code-quality bar on more than 70% of tasks. The talk covers how the validation agent enables reliable evaluation of open-ended feature solutions, how the taste judge evaluates code quality relative to existing repository practices, and what the results reveal about model performance, efficiency, and growing benchmark awareness. 🔗 Links 🌐 Benchmark & Leaderboard: https://senior-swe-bench.snorkel.ai/ 💻 GitHub: https://github.com/snorkel-ai/senior-swe-bench-v2026.06 📖 How Senior SWE-Bench works: https://senior-swe-bench.snorkel.ai/blog/2026-06-16-how-it-works 👤 Henry Ehrenberg: https://snorkel.ai/author/henry-ehrenberg/ 📖 Blog post with full transcript: https://snorkel.ai/blog/senior-swe-bench-evaluating-coding-agents-like-senior-engineers/ 🎙️ Reading Group info: https://snorkel.ai/reading-group/ Previous Reading Group sessions: ▶️ Agents’ Last Exam — https://snorkel.ai/blog/agents-last-exam-reading-group/ ▶️ Olmix: Data Mixing for LLM Development — https://snorkel.ai/blog/olmix-data-mixing-reading-group/ ▶️ Code World Models & AutoHarness — https://snorkel.ai/blog/code-world-models-and-autoharness-for-llm-agents/ ▶️ Collaborative Gym — https://snorkel.ai/blog/collaborative-gym-a-framework-for-enabling-and-evaluating-human-agent-collaboration/ ▶️ JudgmentBench — https://snorkel.ai/blog/judgment-bench-comparing-rubric-and-preference-evaluation-for-quality-assessment-legal-ai/

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close