A chat with the Terminal-Bench team: Stanford & Laude - CLI Agents and Harbor | Snorkel AI Interview
Snorkel AI Chief Scientist Fred Sala and Snorkel Developer Advocate Kobie Crawford had a chance to catch up with Terminal-Bench project leads Mike Merrill, postdoc at Stanford advised by Prof. Ludwig Schmidt, and Alex Shaw, Founding Member of Technical Staff at Laude Institute. Recorded on November 5, 2025, just before the release of Terminal-Bench 2.0 and the Harbor framework, the conversation gave us a peek into the thought process and design decisions underpinning the benchmark, and shared insights about how Harbor supports a wide range of use cases beyond benchmarking. Read the full transcript: https://snorkel.ai/blog/chat-with-the-terminal-bench-team/ Explore Snorkel’s NeurIPS presence and research work: https://snorkel.ai/neurips-event/ More from Snorkel AI Research: https://snorkel.ai/blog/ 0:00 - Introductions 01:27 - Rationale for creating Terminal-Bench 05:53 - Benchmark design considerations 08:25 - Design constraints and impact on agent evaluation 13:03 - Building the Terminal-Bench community 17:55 - Laude Institute and the project philosophy 21:41 - Introducing the Harbor execution framework 25:29 - Transitioning to Terminal-Bench 2 and Harbor 28:56 - What’s exciting about Terminal-Bench 2.0?