Get Started with Open Model Routing | Nemotron Labs
Route smarter across open and closed AI models — NeMo Switchyard + Cognition's Devin Fusion show how agentic workflows cut LLM costs by 60% without sacrificing frontier-level quality. One model is no longer enough for complex AI workflows. The developers getting the most out of agentic AI are building systems of models — routing each step to the right intelligence based on task complexity, cost, latency, and quality. The question isn't which model to use. It's how to route intelligently across all of them. In this session, Joon Hee Lee (Technical Staff Building Devin, Cognition) walks through Devon Fusion, Cognition's hybrid model harness that pairs a frontier lead model with a cheaper sidekick agent — achieving up to 60% cost reduction on the Frontier Code benchmark while maintaining frontier-level performance. Tanay Varshney (Principal Product Research Engineer, NVIDIA) then breaks down the theory and practice of open model routing from first principles, including how NeMo Switchyard — NVIDIA's open-source routing library — enables intelligent routing across any combination of open and closed models in real agent workflows. Together they cover the sidekick architecture (persistent context, prompt caching, task delegation scope), dynamic mid-session routing and cache miss tradeoffs, pre-fill routing using internal LLM state as a confidence signal, and how to pick the right routing approach for your workload — from simple LLM-as-judge all the way to domain-specific engineered harnesses. 0:00 Intro 1:51 Devon Fusion: How Cognition routes coding agents 5:02 Sidekick architecture — lead + sub-agent pair programming 11:20 Dynamic mid-session routing & cache miss tradeoffs 14:15 Meet the team — speakers & Q&A opens 15:56 Can you route between local and frontier endpoints? 25:08 Sidekick vs. frontier model as CTO reviewer? 29:00 Context management for intelligent task delegation 32:48 Open routing & when to switch models 37:25 Is NeMo Switchyard still in alpha? 37:41 Which model actually does the routing? 43:13 Best routing signal: classifier, draft-escalate, or logprob? 44:24 How does a sub-agent avoid context drift? 48:09 Is cache prefill solved? 48:33 Resources — learning paths, GitHub & Discord 50:47 How did Tanay & Jun get into model routing? —————————————————————————— RESOURCES 🔗 NeMo Switchyard GitHub → https://nvda.ws/4xWLigt 📖 NeMo Switchyard tech blog → https://nvda.ws/46lydS6 📖 Cognition Devon Fusion blog → https://cognition.com/blog/devin-fusion 🎓 Agentic AI Learning Paths → https://developer.nvidia.com/topics/ai/agentic-ai-learning-path 💻 Nemotron GitHub → https://github.com/NVIDIA-NeMo/Nemotron 💬 NVIDIA Developer Discord (#nemotron-models) → https://nvda.ws/46Rxucr 📧 [email protected] Speakers: ↳ Joon Hee Lee (Cognition) → https://www.linkedin.com/in/joonheelee ↳ Tanay Varshney (NVIDIA) → https://www.linkedin.com/in/tanayvarshney ↳ Chris Alexiuk (NVIDIA) → https://www.linkedin.com/in/csalexiuk #nvidia #nemotron #aiagents #agenticai #llm #opensource #aiengineering #machinelearning #modelrouting #codingagents #developertools