OSWorld 2.0: Long-Horizon Computer-Use Benchmark | Snorkel AI Reading Group

Snorkel AI
425 views September 3, 2026

At our latest Snorkel AI Reading Group, Mengqi Yuan presents OSWorld 2.0, a benchmark of 108 long-horizon, real-world computer-use workflows spanning 31 self-hosted websites and professional desktop apps. Tasks average more than 300 agent steps and 27 scoring checkpoints, and 69.6% take a skilled human over an hour to complete, a sharp jump from OSWorld 1.0's roughly 30-step tasks. Can today's frontier agents handle the cross-source reasoning, dynamic pop-ups, streaming interaction, and implicit-state inference that real, messy computer-use work demands? The best system in the paper completes only 20.6% of tasks outright, though a Claude model released the day before this talk pushed partial credit above 60%. The talk covers what changed since OSWorld 1.0, the ten challenge phenomena built into the benchmark, four recurring agent failure modes, and the safety checkpoints woven into the task suite, from leaked API keys to destructive file deletions. Chapters 00:00 Introduction 02:59 What is OSWorld, and why OSWorld 1.0 saturated 04:17 Benchmark scale: 300–500 step tasks across multiple apps 05:01 Example task: the multi-app reimbursement workflow 07:50 Letting agents ask a simulated user for clarification 08:39 By the numbers: 31 self-hosted sites, 27 checkpoints per task 10:35 Ten challenge phenomena: cross-source reasoning to implicit state 14:00 Conflicting information: the concert-seat example 15:15 Tutorials, dynamic environments, and streaming tasks 17:40 Live showcase: dynamic pop-ups and streaming interaction 18:22 Self-hosting 31 websites for a stable environment 19:06 How tasks are sourced, designed, and verified 20:43 Frontier agent results: Claude vs. GPT 22:09 Domain-level differences between Claude and GPT 22:54 Four failure modes: information, perception, verification, memory 25:55 Safety checkpoints: leaked API keys and destructive actions 26:53 Q&A 49:52 Wrap-up and the Frontier Data Summit 🔗 Links 📄 Paper: https://arxiv.org/abs/2606.29537 🌐 Benchmark & Leaderboard: https://osworld-v2.xlang.ai/ | https://snorkel.ai/leaderboard/os-world-2-0/ 💻 GitHub: https://github.com/xlang-ai/OSWorld-V2 🐦 Mengqi Yuan on X: https://x.com/yuan_mengq43669 🔗 Mengqi Yuan on LinkedIn: https://www.linkedin.com/in/mengqi-yuan-ab850a323/ 📖 Blog post with full transcript: https://snorkel.ai/blog/osworld-2-0-why-computer-use-agents-fail-most-tasks 🎙️ Reading Group info: https://snorkel.ai/reading-group/ Previous Reading Group sessions: ▶️ Train-to-Test (T²) Scaling Laws: https://snorkel.ai/blog/train-to-test-scaling-laws-reading-group/ ▶️ Senior SWE-Bench: Evaluating Coding Agents: https://snorkel.ai/blog/senior-swe-bench-evaluating-coding-agents-like-senior-engineers/ ▶️ Agents' Last Exam: https://snorkel.ai/blog/agents-last-exam-reading-group/ ▶️ JudgmentBench: Rubric vs. Preference Evaluation: https://snorkel.ai/blog/judgment-bench-comparing-rubric-and-preference-evaluation-for-quality-assessment-legal-ai/ ▶️ Collaborative Gym: Human-Agent Collaboration: https://snorkel.ai/blog/collaborative-gym-a-framework-for-enabling-and-evaluating-human-agent-collaboration/ ▶️ Code World Models & AutoHarness: https://snorkel.ai/blog/code-world-models-and-autoharness-for-llm-agents/ ▶️ Olmix: Data Mixing for LM Development: https://snorkel.ai/blog/olmix-data-mixing-reading-group/

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close