How to Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard

NVIDIA Developer
2,522 views August 11, 2026

AI agents increasingly rely on multiple models for different kinds of work. But the model that’s right for a task can change as an agent works through it. NVIDIA NeMo Switchyard is an open-source library for dynamically routing agent workloads across models based on signals from the request, agent state, tool results, model capabilities, cost, latency and more. In this video, we break down the core components of model routing, including model pools, signals, decision policies, execution and feedback. You’ll also see how different routing approaches work, from LLM classifiers and cascade routing to trainable prefill routers. Learn how model routing can keep routine work on efficient models and move to more capable models when a task demands it, bringing specialized, local and frontier models together in one agentic system. 📝 Tech blog: https://nvda.ws/3Skv2H2 🔗 Get Started: https://github.com/NVIDIA-NeMo/Switchyard 00:00 – The Shift: Why One Model Isn’t Enough 00:28 – A System of Models 01:14 – Introducing NVIDIA NeMo Switchyard 01:28 – The Five Stages of Dynamic Routing 01:43 – Model Pool: Choosing Complementary Models 02:10 – Signals: What the Router Sees 02:45 – Execution: Router, Proxy and Backends 03:14 – Feedback: Quality, Cost, Latency and Reliability 03:39 – Decision Policy: How Switchyard Chooses 03:53 – LLM Classifier Routing 04:39 – Evidence-Driven Cascade 05:08 – Pre-Fill Routing #NVIDIANeMo #ModelRouting

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close