Extreme Co-Design for Efficient Tokenomics and AI at Scale

NVIDIA
4,293 views February 12, 2026

As AI enters the era of real-time reasoning, the key metric for deploying AI at scale is now cost per token — how much it costs to generate intelligence. Reasoning models like mixture-of-experts (MoE) generate massive volumes of tokens to deliver higher-quality results, placing pressure on the entire system — from compute and memory to networking, storage, and software. Featuring insights from NVIDIA, Signal65, Microsoft Azure, and CoreWeave, this discussion explains why extreme co-design — optimizing the full stack as a unified system — is essential to lowering cost per token and maximizing AI ROI, making end-to-end system design the most powerful lever for scaling efficient AI. Learn more: https://blogs.nvidia.com/blog/inference-open-source-models-blackwell-reduce-cost-per-token/

Keyboard shortcuts

On. Switch them off if they collide with your assistive tools; ? still opens this sheet.

Go to

Press g then the letter.

  • gh Latest
  • gs Sources
  • gm Media
  • gv Videos
  • gp Podcasts
  • gc Calendar
  • gd Decoder
  • gz Dataviz
  • ga Datasets
  • gb Blog
  • gk Markets
  • gj Careers
  • gn Prompt Notebook

On this page

  • / Focus search, where there is one
  • t Back to top
  • ? This list
  • Esc Close