Nemotron 3 Ultra eval speedrun: OpenRouter vs Nebius Token Factory

Nebius
241 views September 10, 2026

Same model, same eval, two inference stacks, side by side from start to finish. NVIDIA released Nemotron 3 Ultra this summer with open weights and a one million token context window. We put it on Token Factory on day zero. To show what that inference stack does, Rob Spectre runs IFBench, Allen AI's instruction-following eval, on the same model twice: once served vanilla through OpenRouter, once on Nebius Token Factory, which builds a customized stack per model with a speculative decoder, KV cache optimizations, and custom CUDA kernels. If you want to build like this, come check out the Nebius AI Builder Program. It bundles credits for Nebius Token Factory, Tavily plus training, office hours, and access to the builder community. Register once at dev.nebius.com/builders. Links: Builder Program: https://dev.nebius.com/builders Token Factory: https://dev.nebius.com/token-factory Chapters 0:00 - What's in the Builder Program 0:40 - Nemotron 3 Ultra from NVIDIA 1:11 - What IFBench measures 1:35 - Setting up OpenBench 2:07 - The speedrun and the inference stack 3:46 - How it ended and how to sign up

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close